Examples with real data¶
Each script in examples/ downloads a real government data release and crosswalks it. Each one shows a different situation: a reused name, renames carried by content, the threshold tradeoff, and format-independence.
pip install "field-match[optimal]"
python examples/cdc_atsdr_svi.py
python examples/cdc_places.py
python examples/imls_pls.py
python examples/namcs_formats.py
Each script saves its files in examples/data/ on first run; drag the same pair into the web app to see the comparison visually. The SVI example is also available as a notebook, for a step-by-step walkthrough instead of a script.
SVI: the reused-name trap¶
CDC/ATSDR Social Vulnerability Index, 2010 vs. 2022 county files. As a script, cdc_atsdr_svi.py; as a notebook, cdc_atsdr_svi.ipynb (open in Colab, no install required).
- Most of the 103 columns renamed between releases.
- Two columns reused with different meanings: in 2010,
STheld text state abbreviations andSTATEheld numeric codes; by 2022 the numeric code is calledST. - A name-based merge would silently corrupt the state column.
- The content check catches it: both columns reported as suspect, true counterparts found instead (
STtoST_ABBR,STATEtoST).
CDC PLACES: content carries the match¶
cdc_places.py: the 500 Cities project becoming PLACES in 2020.
- Renamed columns are flattened lowercase words:
citynamevs.locationnamescores 0 on name similarity. - Values carry the whole match:
citynametolocationname,cityfipstolocationid,populationcounttototalpopulation. - API technique demonstrated: use the reference file's own FIPS codes to request the comparable slice from the new data source, so contents genuinely overlap.
IMLS PLS: the threshold tradeoff¶
imls_pls.py: thirty years of the Public Libraries Survey, 1992 vs. 2022.
- That much drift thins content overlap; several true renames (
LIB_NAMEtoLIBNAME,LIB_ADDRtoADDRESS) score just under the defaultmatch_threshold. - Lowering it to 0.4 recovers them, along with one wrong suggestion that review catches: the threshold tradeoff in miniature.
NAMCS: three formats, one answer¶
namcs_formats.py: National Ambulatory Medical Care Survey, 2022 vs. 2024, read three ways, once each from the SAS, Stata, and R releases NCHS publishes.
- All three give byte-for-byte identical results: 63 verified, 12 suspect (the
DX1-DX11diagnosis codes and theVISWTsurvey weight, both genuinely drifted as the sample nearly doubled), 0 renamed, dropped, or added. - The format the data arrived in does not change the answer.
Note: this script downloads all six files, about 590 MB total. Trim its FORMATS list to just the format you use.