Producing Research-Ready Restricted Use Decennial Census Microdata for Puerto Rico and Island Area, 1970, 1980, 1990

John Sullivan, Census Bureau
Gina A. Erickson, University of Minnesota
Todd K. Gardner, U.S. Census Bureau

The U.S. Constitution mandates that the Census Bureau conduct a census of the population every ten years. The goals of the Decennial Census Digitization and Linkage Project (DCDL) are to fill a gap in digitized data and the longitudinal data infrastructure by creating and linking research-ready data from the censuses and surveys conducted by the US Census Bureau. This includes the data from US territories (commonly referred to as Island Areas or Outlying Areas by the Census Bureau) and Puerto Rico. Puerto Rico has been enumerated in each Decennial Census since 1910 with U.S. Territories and Commonwealths following in 1920 (Guam and American Samoa), 1940 (U.S. Virgin Islands) and 1970 (Commonwealth of the Northern Mariana Islands). However, research-ready restricted use decennial microdata had not been produced for Puerto Rico and the Island Areas for 1970, 1980 and 1990. Decades after data collection, we located the unlabeled data and made efforts to make it available for research. Difficulty in locating appropriate documentation and differences between Puerto Rico, the Island Areas and the 50 states in enumeration procedures and instruments created challenges for determining record layouts and harmonizing data fields. This paper describes the location of the raw decennial data, the use of varied documentation to create record layouts, efforts to harmonize data with existing restricted decennial microdata and quality checks on the production files. The research-ready data from both the sample edited file “long-form” and high-density file “short-form” add 3-4 million person records in each decennial year to the hundreds of millions of records that will soon be included in the Census Bureau’s longitudinal linkage infrastructure through the DCDL project.

No extended abstract or paper available

 Presented in Session 14. Using Data and Language Models