patient_identity
Patient identity: deriving a BitfountPatientID, and folding scans onto it.
One derivation, shared by every step route that has name parts rather than a
ready-made ID: ehr_query (which reads them off an EHR patient record) and
patient_enrichment (which reads them off a supplied file). Both wrap
generate_bitfount_patient_id, so a patient reaching us down either route —
or down the imaging route that hashes a DICOM name directly — keys the same
ID. A second derivation that drifted from it would split one patient into
two across the caches that key on this value.
ScanIdentitySource folds the scan rows onto that derivation, giving one
identity and one acquisition time per file. It lives here rather than beside
the wave planner that first needed it because the reduce step needs the same
fold, and a step must not import the flows layer.
Module
Functions
bitfount_patient_id_from_name_dob
def bitfount_patient_id_from_name_dob( name: str | None, dob: str | date | datetime | None, *, dayfirst: bool | None = None,) ‑> str | None:Generate a BitfountPatientID from a full name and date of birth.
Wraps generate_bitfount_patient_id on a single-row DataFrame so every
name-and-DOB route produces IDs identical to the imaging-file route.
A date of birth that will not parse returns None: the underlying
_identity_date_key yields pd.NA and the derivation propagates it, so a
dob of "N/A" or "unknown" mints no ID rather than a well-formed one
keyed on that literal string, which would join to nobody while being
indistinguishable from a real one.
A date that parses but whose field order is ambiguous is not rejected
either, and is a sharper problem: 04/03/1952 read month-first keys a
different ID from the same patient's 1952-03-04. One row cannot settle
its own order, so a caller holding the whole column resolves it once with
resolve_date_order and passes the answer; a caller with only this row
gets the configured order, and the month-first default is reported once
per batch rather than silently applied.
Arguments
name: The full name, in any formsplit_full_nameaccepts — typically the caret formfull_name_from_partsbuilds.dob: The date of birth, in any formpd.to_datetimeaccepts.dayfirst: Read a two-field date day-first (True) or month-first (False).Noneresolves it frompatient_id_date_order, which cannot infer from a single row.
Returns
The ID, or None when the name is missing, the name has no given
and family pair to key on, the date of birth is missing, or
generation fails.
full_name_from_parts
def full_name_from_parts(family_name: str | None, given_name: str | None) ‑> str | None:Build a full-name string from family/given name parts.
The parts are joined in the DICOM Family^Given form rather than with a
space. Everything that consumes this string re-parses it with
split_full_name, and a space-separated name is ambiguous there: it is
read as Given Family, so "Doe Jane" would come back as given "Doe"
and family "Jane". The caret form states which component is which, so
the parts survive the round trip and this route derives the same
BitfountPatientID as the imaging-file route does for the same patient.
Arguments
family_name: The family name, if the source states one.given_name: The given name, if the source states one.
Returns
The caret-joined name, or None when neither part is available.
Classes
FileIdentity
class FileIdentity( bitfount_patient_id: str | None, scan_datetime: datetime | None, patient_key: str | None = None,):What the scan rows say about one file.
Attributes
bitfount_patient_id: The derived patient id, orNonewhen the file's rows carry no usable name and date of birth.scan_datetime: The file's most recent acquisition time, orNone.patient_key: The source's own patient identifier, orNonewhen the file's rows state none. Carried, not derived from: it is per-source and so cannot key a patient across devices, but it distinguishes two patients the derived id collapses.
Variables
- static
bitfount_patient_id : str | None
- static
patient_key : str | None
- static
scan_datetime : datetime.datetime | None
ScanIdentitySource
class ScanIdentitySource(rows: Iterable[ScanIdentity]):Per-file patient identity and acquisition time, from the scan rows.
Built once per run from a single streamed pass over scan_metadata. A file
with several series contributes several rows; the identity is taken from the
first that yields one and the recency from the newest across them all.
Identity comes from ScanIdentity.bitfount_patient_id, the one shared
derivation, so a patient waved here keys the same as the patient the
eligibility steps roll up — a second derivation that drifted would split one
patient into two and put their scans in different waves.
Held per file rather than per row: three fields and the path key, so a
200k-file datasource costs tens of MB, and only for as long as
WavePlanner.candidates takes to join the inventory against it.
Arguments
rows: The scan rows, asiter_scan_identitiesyields them.
Variables
- static
rows : Iterable[ScanIdentity]
Methods
for_file
def for_file(self, file_path: str) ‑> FileIdentity | None:Return what the scan rows say about file_path.
Arguments
file_path: The file to look up.
Returns
Its identity, or None when the scan sweep has produced no row for
it yet. Not a permanent state: the sweep selects files the same way
the run's own inventory does (select_for_datasource), and writes a
reason row for a file it could not parse, so every file here ends
up with a row once the sweep reaches it.