Old Text, New Tricks? Combining AI and Established Methods for Historical Document Analysis

Daniel Moulton, Federal Reserve Bank of Philadelphia
Robyn Smith, Federal Reserve Bank of Philadelphia
Larry Santucci, Federal Reserve Bank of Philadelphia

For thousands of years administrative processes have generated unstructured textual data. Twenty-five years into the 21st century much of this data remains locked up in original physical form, as simple images, or in digital formats that limit analysis. Advances in Optical Character Recognition (OCR), full-text search (OpenSearch), and Large Language Models (LLMs) have made the process of extracting, processing, and analyzing this data at scale more accessible than ever before. Focusing on 55 years of Philadelphia property deeds and the identification of racially restrictive covenants, we examine a dataset that is highly homogeneous yet presents varying challenges through different document formats – handwritten, typewritten, and photostat – exhibiting a wide range of microfilm and scan quality. We provide a systematic assessment of the accuracy-cost tradeoffs across different OCR, search, and LLM solutions. The paper examines how LLMs can be leveraged for complex information extraction and structuring tasks while maintaining accuracy and cost efficiency. Importantly, we demonstrate that LLMs are not necessarily superior to alternative solutions, showing where more established, deterministic approaches such as full-text search may be more appropriate. We argue that this case study offers lessons that generalize to a wide class of social research questions. Our findings provide a methodological roadmap for scholars seeking to leverage administrative textual data at scale, while being mindful of both the possibilities and limitations of mature versus emerging technologies. The views expressed here are solely those of the authors and do not necessarily reflect the views of the Federal Reserve Bank of Philadelphia or the Federal Reserve System.

No extended abstract or paper available

 Presented in Session 107. AI Impact on Data Infrastructure