COVID-19 Records: NLP Pipeline
A Python pipeline that turns COVID-19 encounter notes into structured records for exploring medications, comorbidities, and demographic trends.
Making encounter notes queryable
A medication can appear under different names across a set of notes. A diagnosis may sit inside a paragraph of observations. Before those records can support comparison, the relevant details need a consistent structure.
I built this pipeline to connect text extraction with dataset analysis. Its stages handle different parts of that problem: identifying information in a note, checking the resulting record, reconciling medication references, and summarizing patterns across records.
From free text to structured records
- Extract candidate fields. A large language model reads the encounter text and produces structured data for the next stage.
- Check the record shape. Pydantic validates the output against defined field types and schema requirements, giving downstream code a consistent format to work with.
- Find similar medication references. FAISS searches vector representations of medication text to help reconcile wording that exact text matching would miss.
Analyzing patterns across records
PySpark supports aggregation across the resulting dataset, including comorbidity patterns, medication frequencies, and demographic trends. Parquet provides columnar storage for the structured output, while visualizations make the aggregate results easier to examine.
The useful change is that an analyst can work with comparable fields across many notes, then examine the patterns those fields reveal.
What validation establishes
Schema validation checks whether an extracted value has the expected structure and type. Clinical accuracy still depends on whether that value faithfully represents the original note. Likewise, a close semantic match identifies a medication candidate; similarity alone cannot establish that two references are equivalent. Those distinctions matter when interpreting the analysis.