refresh
Prefect flow for periodic re-indexing of stale file metadata.
This module implements the cron-driven counterpart to file_metadata_runtime.
While the initial indexing is triggered once by the pod when it starts,
metadata rows grow stale over time as files are added, modified or removed.
file_metadata_refresh re-indexes any datasource that has not completed an
indexing pass within a configurable staleness window.
Design
A single long-lived flow.serve() call is started in a daemon thread at
orchestrator startup. It registers one deployment on the embedded Prefect
server and polls for scheduled work according to the configured cron expression.
Because there is only one deployment (not one per datasource), no lifecycle
management is required when pods are created or deleted.
At each scheduled tick the flow:
- Scans
~/.bitfount/cache/pods/for pod directories that contain abackground_cache.dbfile. - For each DB, enumerates the distinct
(pod_name, datasource_name, datasource_type, task_hash)groups inFileMetadataand keeps those whose last completedfile_metadatarun finished more than staleness_days ago, or that have none. - Drops any datasource the pod has not recorded as configured, and skips a
pod DB altogether when it holds no such record. See "Removed datasources"
below. Then drops any datasource the pod last recorded as linked only to
archived projects, and orders the rest by the pod's last health check,
then never-indexed first, the Hub's
configurationLastUpdatednewest first, and least recently indexed. See "Indexing order" below. - Skips any datasource for which a live (non-terminal) Prefect flow-run is
already tagged with its
task_hash— this prevents a second indexing pass from starting while the previous one (or the pod's start-up run) is still in progress._is_run_activeis still called per datasource, but only for its side effect of reaping orphaned"running"bookkeeping rows (decision: Prefect owns liveness/dedup, therunstable is bookkeeping-only cleanup, not a gate). - Submits a
background/file_metadata_runtimedeployment run for each stale datasource viarun_deployment(..., timeout=0), tagged withtask_hash:<task_hash>andrun_type:file_metadataso future dedup checks (here and at the orchestrator's initial trigger) see it. Using the deployment path (rather than calling the flow directly as a sub-flow) ensures the same concurrency limits and deployment-level configuration that apply to orchestrator-triggered runs also apply here. Runs are parked in submission order and started as slots free (seebitfount.runtimes.queueing), so they run in that order.
Input reconstruction
Because the provenance columns (pod_name, datasource_name,
datasource_type, source_path) were written to every FileMetadata
at index time, the refresh flow reconstructs a re-index's inputs without
re-reading the pod config JSON. The original input path or connection string
is read directly from the source_path column; the legacy
os.path.commonpath derivation is only used for rows indexed before that
column was added.
Removed datasources
Those same columns say nothing about whether the datasource still exists.
FileMetadata rows outlive a datasource's removal from the pod config — they
are shared inventory keyed by file path, so they are deliberately not deleted
with it — and enumerating stale datasources from them alone re-indexes a
datasource that was dropped weeks ago, at whatever size it was when it was
dropped.
This flow runs in the orchestrator and has no pod handle, so it reads the
pod's own answer instead: the configured_datasources snapshot the pod writes
at startup and on every config reload. A datasource with no live row in it is
not re-indexed, and a pod DB with no rows at all is skipped entirely rather
than assumed configured — a cache that predates the snapshot cannot
distinguish the two cases, and the pod rewrites it on its next start.
Indexing order
The pod records, with that snapshot, what it last knew about each datasource:
its health, the Hub's configurationLastUpdated, and whether every project it
is linked to is archived. When each was last indexed is read from runs.
This flow cannot ask the Hub, so it orders and skips by that record, with the
rule the pod uses at start-up (runtimes.file_metadata.prioritisation). A
datasource with no record is indexed, after those that have one.
Module
Functions
file_metadata_refresh
async def file_metadata_refresh(staleness_days: int = 1) ‑> int:Re-index stale file metadata across all pods.
Scans every pod's background_cache.db under
~/.bitfount/cache/pods/, identifies datasources not indexed within
staleness_days, and submits a background/file_metadata_runtime
deployment run for each one — unless a non-stale indexing run for that
datasource is already active.
This flow is intended to be registered as a single long-lived Prefect
deployment via file_metadata_refresh.serve(cron=...) on orchestrator
startup. It uses global_limit=1 so only one refresh pass runs at a
time even if a scheduled tick fires while the previous pass is still
running.
Runs are submitted as fire-and-forget (timeout=0) so the refresh flow
returns as soon as all runs are scheduled, not when they complete.
is_flow_run_active (a Prefect flow-run tag query) prevents duplicate
runs from being scheduled for the same datasource; _is_run_active is
still called per datasource but only reaps orphaned bookkeeping rows.
Arguments
staleness_days: Number of days after which aFileMetadatais considered stale and eligible for re-indexing. Defaults to 1. Staleness is measured from the datasource's last completedfile_metadatarun, so it is not stale again until this window elapses — this value, not the deployment's cron, is what sets the re-walk cadence. Ticking more often than the window does nothing.
Returns
The total number of file_metadata_runtime runs triggered.