Skip to main content

refresh

Prefect flow for periodic re-indexing of stale file metadata.

This module implements the cron-driven counterpart to file_metadata_runtime. While the initial indexing is triggered once by the pod when it starts, metadata rows grow stale over time as files are added, modified or removed. file_metadata_refresh re-indexes any datasource that has not completed an indexing pass within a configurable staleness window.

Design​

A single long-lived flow.serve() call is started in a daemon thread at orchestrator startup. It registers one deployment on the embedded Prefect server and polls for scheduled work according to the configured cron expression. Because there is only one deployment (not one per datasource), no lifecycle management is required when pods are created or deleted.

At each scheduled tick the flow:

  1. Scans ~/.bitfount/cache/pods/ for pod directories that contain a background_cache.db file.
  2. For each DB, enumerates the distinct (pod_name, datasource_name, datasource_type, task_hash) groups in FileMetadata and keeps those whose last completed file_metadata run finished more than staleness_days ago, or that have none.
  3. Drops any datasource the pod has not recorded as configured, and skips a pod DB altogether when it holds no such record. See "Removed datasources" below. Then drops any datasource the pod last recorded as linked only to archived projects, and orders the rest by the pod's last health check, then never-indexed first, the Hub's configurationLastUpdated newest first, and least recently indexed. See "Indexing order" below.
  4. Skips any datasource for which a live (non-terminal) Prefect flow-run is already tagged with its task_hash — this prevents a second indexing pass from starting while the previous one (or the pod's start-up run) is still in progress. _is_run_active is still called per datasource, but only for its side effect of reaping orphaned "running" bookkeeping rows (decision: Prefect owns liveness/dedup, the runs table is bookkeeping-only cleanup, not a gate).
  5. Submits a background/file_metadata_runtime deployment run for each stale datasource via run_deployment(..., timeout=0), tagged with task_hash:<task_hash> and run_type:file_metadata so future dedup checks (here and at the orchestrator's initial trigger) see it. Using the deployment path (rather than calling the flow directly as a sub-flow) ensures the same concurrency limits and deployment-level configuration that apply to orchestrator-triggered runs also apply here. Runs are parked in submission order and started as slots free (see bitfount.runtimes.queueing), so they run in that order.

Input reconstruction​

Because the provenance columns (pod_name, datasource_name, datasource_type, source_path) were written to every FileMetadata at index time, the refresh flow reconstructs a re-index's inputs without re-reading the pod config JSON. The original input path or connection string is read directly from the source_path column; the legacy os.path.commonpath derivation is only used for rows indexed before that column was added.

Removed datasources​

Those same columns say nothing about whether the datasource still exists. FileMetadata rows outlive a datasource's removal from the pod config — they are shared inventory keyed by file path, so they are deliberately not deleted with it — and enumerating stale datasources from them alone re-indexes a datasource that was dropped weeks ago, at whatever size it was when it was dropped.

This flow runs in the orchestrator and has no pod handle, so it reads the pod's own answer instead: the configured_datasources snapshot the pod writes at startup and on every config reload. A datasource with no live row in it is not re-indexed, and a pod DB with no rows at all is skipped entirely rather than assumed configured — a cache that predates the snapshot cannot distinguish the two cases, and the pod rewrites it on its next start.

Indexing order​

The pod records, with that snapshot, what it last knew about each datasource: its health, the Hub's configurationLastUpdated, and whether every project it is linked to is archived. When each was last indexed is read from runs. This flow cannot ask the Hub, so it orders and skips by that record, with the rule the pod uses at start-up (runtimes.file_metadata.prioritisation). A datasource with no record is indexed, after those that have one.

Module​

Functions​

file_metadata_refresh​

async def file_metadata_refresh(staleness_days: int = 1) ‑> int:

Re-index stale file metadata across all pods.

Scans every pod's background_cache.db under ~/.bitfount/cache/pods/, identifies datasources not indexed within staleness_days, and submits a background/file_metadata_runtime deployment run for each one — unless a non-stale indexing run for that datasource is already active.

This flow is intended to be registered as a single long-lived Prefect deployment via file_metadata_refresh.serve(cron=...) on orchestrator startup. It uses global_limit=1 so only one refresh pass runs at a time even if a scheduled tick fires while the previous pass is still running.

Runs are submitted as fire-and-forget (timeout=0) so the refresh flow returns as soon as all runs are scheduled, not when they complete. is_flow_run_active (a Prefect flow-run tag query) prevents duplicate runs from being scheduled for the same datasource; _is_run_active is still called per datasource but only reaps orphaned bookkeeping rows.

Arguments

  • staleness_days: Number of days after which a FileMetadata is considered stale and eligible for re-indexing. Defaults to 1. Staleness is measured from the datasource's last completed file_metadata run, so it is not stale again until this window elapses — this value, not the deployment's cron, is what sets the re-walk cadence. Ticking more often than the window does nothing.

Returns The total number of file_metadata_runtime runs triggered.