repository
Cache-backed implementation of the patient-data API's data-access seam.
Provides CacheBackedPatientDataRepository, which serves the API from the
pod's background cache tables. The PatientDataRepository Protocol it
implements lives in protocol.py.
Module
Functions
build_evidence_mapping
def build_evidence_mapping(evidence: Mapping[str, Any]) ‑> dict[str, str]:Map criteria-matching config input fields to their evidence criterion key.
Rebuilt from the evidence itself: each criterion in evidence["criteria"]
carries the config input field(s) that produced it (config_fields, a
{config_field: {operator: operand}} map stamped at evidence-build time),
so this inverts its keys to {config_field: output_field} without consulting
any static config→criterion table. Lets a caller resolve "which evidence
entry relates to total_ga_area_lower_bound?" directly. A range criterion's
lower + upper bound both resolve to the one merged criterion.
Arguments
evidence: Apatient_eligibilityevidence object (with acriteriasub-map keyed byoutput_field), or any mapping carryingcriteria.
Returns
A {config_field: output_field} map over the criteria present. Empty for
evidence lacking config_fields (e.g. rows written before it was added).
Classes
CacheBackedPatientDataRepository
class CacheBackedPatientDataRepository( cache_getter: Callable[[], CacheProtocol], image_service: ScanImageService | None = None,):PatientDataRepository backed by the pod background cache.
Resolves each project_id to its authoritative partitions via the
trials_published_data_pointer pointer, so every trial's data comes from a
single published run and partitions are never mixed. The lookup is always
paired with the table being read (_published_partitions): each step writes
under its own Merkle hash, so there is no one hash that serves every table.
Arguments
cache_getter: A zero-arg callable returning the pod background cache. Called lazily and memoised on first use.image_service: The scan-image renderer this repository dispatches to once a scan reference is resolved.None(the default) leaves the scan-image feature disabled, soget_scan_imagereports it as not found rather than calling a method onNone.
Methods
get_eligibility_counts
def get_eligibility_counts( self, trial_ids: list[str],) ‑> dict[str, EligibilityCounts]:Return per-trial status counts, zeroed for unknown trials.
Arguments
trial_ids: The trial (project) IDs to count.
Returns
A map of trial ID → EligibilityCounts.
get_eligibility_evidence
def get_eligibility_evidence( self, bitfount_patient_id: str, project_id: str,) ‑> EligibilityEvidence | None:Return rolled-up evidence for a (patient, trial), or None.
Arguments
bitfount_patient_id: The Bitfount patient ID.project_id: The trial (project) ID.
Returns
The EligibilityEvidence, or None if there is no rollup.
get_patient
def get_patient( self, bitfount_patient_id: str, project_id: str | None = None,) ‑> PatientSummary | None:Return a single patient summary, or None.
Resolves the identity row and eligible-trial set with targeted primary-key lookups rather than materialising every patient.
Arguments
bitfount_patient_id: The Bitfount patient ID.project_id: When given, scopeslast_analysed_image_dateto that trial; omitted, it spans every trial the pod publishes.
Returns
The PatientSummary, or None if the patient is unknown.
get_scan_image
def get_scan_image( self, *, scan_id: str | None, path: str | None, modality: str | None, laterality: str | None, width: int | None, refresh: bool, include_masks: bool = False, include_vectors: bool = False, segmentation_classes: list[str] | None = None,) ‑> bytes:Resolve the scan reference and dispatch to the image service.
Arguments
scan_id: The scan identifier, or None (exactly one of scan_id/path).path: An explicit, allowlisted file path, or None.modality: Optional series selector for multi-series files.laterality: Optional "L"/"R" series selector.width: Optional output width (px); None uses the configured default.refresh: When True, bypass and overwrite any cached archive.include_masks: Whether the archive should carry a rendered mask PNG per segmentation class; passed straight through to the image service.include_vectors: Whether the served index should carry each class's raw drawable instances; also passed straight through.segmentation_classes: Optional class-name filter, applied only when an overlay output was requested.
Returns
The application/zip archive bytes.
Raises
ScanNotFoundError: The scan-image feature is not enabled, or the reference does not resolve to an authorised file.ScanRequestError: Not exactly one ofscan_id/pathwas supplied, or the width/file is invalid (raised by the image service).
list_patients
def list_patients( self, *, search: str | None = None, trials: list[str] | None = None, statuses: Collection[EligibilityStatus] | None = None, project_id: str | None = None, page: int = 1, page_size: int = 50,) ‑> PatientList:Return one page of patients, filtered and searched in the database.
Filtering, search and windowing run in SQL (name/ehr_patient_id are
plaintext columns), so only the requested page is materialised. Each
returned patient keeps its full trial_statuses regardless of the
trials/statuses filter, so a listing narrowed to one trial still
reports each patient's verdicts across every trial the pod publishes.
Arguments
search: Case-insensitive substring matched against name or EHR id.trials: If given, restrict to patients with a verdict for any of these trials (union). Blank entries are dropped, so a filter of nothing but blanks is no filter at all and serves every patient; a filter naming only trials with no published partition does restrict, to nobody, and yields an empty page.statuses: Which verdictstrialsadmits.Noneadmits all three, so the filter selects those trials' whole cohorts — every patient evaluated, however it came out. Ignored whentrialsisNone, since there is no trial to read a verdict under.project_id: When given,last_analysed_image_dateis the most recent scan-processing time for that trial. Omitted, it is the most recent across every trial the pod publishes, which next to trial-scoped neighbours reads as a contradiction — a trial-aware caller should pass it.page: 1-indexed page number.page_size: Page size (number of items per page).
Returns
A PatientList whose items are the requested page and whose
total is the count of all matching patients before pagination.
list_scan_evidence
def list_scan_evidence( self, bitfount_patient_id: str, project_id: str, *, page: int = 1, page_size: int = 50,) ‑> ScanEvidenceList:Return a patient's per-scan evidence for a trial, one page at a time.
Scoped to the trial's published scan_eligibility partition, narrowed
from the all-projects map so a scan analysed under a different trial
cannot leak in.
Sorted most recent study_date first with scan_id breaking ties —
matching how roll_up_patient picks its determining scan, so the scan
the rollup names tends to sort near the top and the order is stable
across calls. Rows are sorted before the page is cut, so paging is
consistent.
Arguments
bitfount_patient_id: The Bitfount patient ID.project_id: The trial (project) ID.page: 1-indexed page number.page_size: Items per page.
Returns
A ScanEvidenceList; empty when the trial publishes no
scan_eligibility partition or the patient has no scans in it.
resolve_scan_path
def resolve_scan_path(self, *, scan_id: str | None, path: str | None) ‑> str | None:Resolve a scan reference to a validated absolute file path.
Arguments
scan_id: The scan identifier to resolve, or None.path: An explicit filesystem path to validate, or None.
Both branches resolve only against the published scan partitions, so a scan whose only rows belong to a superseded run is not reachable — the same scoping the eligibility reads use.
Returns The validated real path, or None if unknown/unauthorised.
Raises
ValueError: If not exactly one ofscan_id/pathis supplied.