Skip to main content

patient_identity

Patient identity: deriving a BitfountPatientID, and folding scans onto it.

One derivation, shared by every step route that has name parts rather than a ready-made ID: ehr_query (which reads them off an EHR patient record) and patient_enrichment (which reads them off a supplied file). Both wrap generate_bitfount_patient_id, so a patient reaching us down either route — or down the imaging route that hashes a DICOM name directly — keys the same ID. A second derivation that drifted from it would split one patient into two across the caches that key on this value.

ScanIdentitySource folds the scan rows onto that derivation, giving one identity and one acquisition time per file. It lives here rather than beside the wave planner that first needed it because the reduce step needs the same fold, and a step must not import the flows layer.

Module​

Functions​

bitfount_patient_id_from_name_dob​

def bitfount_patient_id_from_name_dob(    name: str | None, dob: str | date | datetime | None, *, dayfirst: bool | None = None,) ‑> str | None:

Generate a BitfountPatientID from a full name and date of birth.

Wraps generate_bitfount_patient_id on a single-row DataFrame so every name-and-DOB route produces IDs identical to the imaging-file route.

A date of birth that will not parse returns None: the underlying _identity_date_key yields pd.NA and the derivation propagates it, so a dob of "N/A" or "unknown" mints no ID rather than a well-formed one keyed on that literal string, which would join to nobody while being indistinguishable from a real one.

A date that parses but whose field order is ambiguous is not rejected either, and is a sharper problem: 04/03/1952 read month-first keys a different ID from the same patient's 1952-03-04. One row cannot settle its own order, so a caller holding the whole column resolves it once with resolve_date_order and passes the answer; a caller with only this row gets the configured order, and the month-first default is reported once per batch rather than silently applied.

Arguments

  • name: The full name, in any form split_full_name accepts — typically the caret form full_name_from_parts builds.
  • dob: The date of birth, in any form pd.to_datetime accepts.
  • dayfirst: Read a two-field date day-first (True) or month-first (False). None resolves it from patient_id_date_order, which cannot infer from a single row.

Returns The ID, or None when the name is missing, the name has no given and family pair to key on, the date of birth is missing, or generation fails.

full_name_from_parts​

def full_name_from_parts(family_name: str | None, given_name: str | None) ‑> str | None:

Build a full-name string from family/given name parts.

The parts are joined in the DICOM Family^Given form rather than with a space. Everything that consumes this string re-parses it with split_full_name, and a space-separated name is ambiguous there: it is read as Given Family, so "Doe Jane" would come back as given "Doe" and family "Jane". The caret form states which component is which, so the parts survive the round trip and this route derives the same BitfountPatientID as the imaging-file route does for the same patient.

Arguments

  • family_name: The family name, if the source states one.
  • given_name: The given name, if the source states one.

Returns The caret-joined name, or None when neither part is available.

Classes​

FileIdentity​

class FileIdentity(    bitfount_patient_id: str | None,    scan_datetime: datetime | None,    patient_key: str | None = None,):

What the scan rows say about one file.

Attributes

  • bitfount_patient_id: The derived patient id, or None when the file's rows carry no usable name and date of birth.
  • scan_datetime: The file's most recent acquisition time, or None.
  • patient_key: The source's own patient identifier, or None when the file's rows state none. Carried, not derived from: it is per-source and so cannot key a patient across devices, but it distinguishes two patients the derived id collapses.

Variables​

  • static bitfount_patient_id : str | None
  • static patient_key : str | None

ScanIdentitySource​

class ScanIdentitySource(rows: Iterable[ScanIdentity]):

Per-file patient identity and acquisition time, from the scan rows.

Built once per run from a single streamed pass over scan_metadata. A file with several series contributes several rows; the identity is taken from the first that yields one and the recency from the newest across them all.

Identity comes from ScanIdentity.bitfount_patient_id, the one shared derivation, so a patient waved here keys the same as the patient the eligibility steps roll up — a second derivation that drifted would split one patient into two and put their scans in different waves.

Held per file rather than per row: three fields and the path key, so a 200k-file datasource costs tens of MB, and only for as long as WavePlanner.candidates takes to join the inventory against it.

Arguments

  • rows: The scan rows, as iter_scan_identities yields them.

Variables​

  • static rows : Iterable[ScanIdentity]

Methods​


for_file​

def for_file(self, file_path: str) ‑> FileIdentity | None:

Return what the scan rows say about file_path.

Arguments

  • file_path: The file to look up.

Returns Its identity, or None when the scan sweep has produced no row for it yet. Not a permanent state: the sweep selects files the same way the run's own inventory does (select_for_datasource), and writes a reason row for a file it could not parse, so every file here ends up with a row once the sweep reaches it.