|
|
Christian Møller Dahl, University of Southern Denmark
Sam Hwang, University of British Columbia
Torben Johansen, University of Southern Denmark
Munir Squires, University of British Columbia
Historical census records enable researchers to track individual outcomes over time, but linking individuals across census rounds is particularly challenging for minority and immigrant populations due to transcription errors in handwritten names. We develop a machine learning approach that improves name transcription in historical U.S. census records, addressing the specific challenges of transcribing unfamiliar names and dense tabular formats. Independent human transcribers disagree on names in 30 percent of records, with higher disagreement rates for foreign-born individuals and non-English speakers. Our machine transcriptions increase linking rates by 147 percent for records where human transcribers disagree, while simultaneously improving match quality by 38 percent. These improvements help expand sample sizes for traditionally under-linked groups - including the foreign-born, non-white residents, and those with no formal schooling - where each additional linked record is particularly valuable for statistical inference. Validation against independent genealogical records confirms these gains represent genuine accuracy improvements rather than spurious matches. Our findings demonstrate that improved transcription methods can substantially enhance research on historically underrepresented populations in linked census data.
Presented in Session 143. Digitizing Analog History