|
|
Caglar Koylu, University of Iowa
Jonas Helgertz, University of Minnesota/Lund University
Evan Roberts, University of Minnesota
Alice Bee Kasakoff, University of South Carolina
Population-scale family trees derived from crowd-sourced genealogy platforms are potentially a transformative resource for historical demographic research, providing multi-generational insights into kinship, migration, and social structure. Yet, the representativeness of family tree data remains largely unexamined. In this paper, we introduce a novel population-scale family tree dataset derived from crowd-sourced genealogical records, featuring a largest connected component of approximately 40 million individuals spanning several centuries. To assess the demographic coverage and inherent biases of this extensive dataset, we link these family tree records with the 1880 U.S. Census using a machine learning algorithm developed within the Multigenerational Longitudinal Panel (MLP) project. This algorithm leverages an extensive set of individual, familial, and contextual characteristics—including first and last names, birth years and places, and kinship ties such as father, mother, spouse, and sibling information—to establish high-confidence matches between tree records and census individuals. Linking tree records with census data not only enhances family trees by incorporating information such as occupation, household composition, and township-level residence but also provides a crucial step toward reconstructing multi-generational lineages that the census alone does not capture. Initial results indicate that our linked dataset represents roughly 3% of the 1880 U.S. population and is predominantly comprised of U.S.-born European descendants. Moreover, our analysis reveals striking geographic variation in representativeness—regions such as Utah exhibit significantly higher linkage rates, reflecting robust localized genealogical practices—while groups including Native Americans, African Americans, Mexicans, and individuals from eastern and southern Europe as well as Ireland remain significantly underrepresented. These findings underscore both the promise of family tree data for historical demographic analysis and the need for further methodological refinements to improve representativeness in large-scale linkage efforts.
No extended abstract or paper available
Presented in Session 195. Linking and Data Quality