“Tables Are Tricky”. Testing Text Encoding Initiative (Tei) Guidelines for Fair Upcycling of Digitised Historical Statistics.

Gabi Wuethrich, University of Zurich

In the context of the digitisation of 1918 pandemic data in Zurich, a project on digital data management tests the implementation of XML structures based on the Text Encoding Initiative (TEI), an XML-based standardized vocabulary for text structures, on historical statistical tables. The basic idea of the data management project was to prepare tables of historical health statistics in a sustainable way to make them reusable, interoperable, and machine-readable in a platform-independent way. After Zurich’s Central Library had retro-digitized and published the serial statistical print publications from the 1910s and 1920s, the goal was to capture their content semi-automatically with OCR in Excel and to prepare them as XML documents according to the TEI guidelines including a fitting XML schema. Such clearly structured XML documents should be relatively easily convertible into formats readable by a variety of statistical tools. Converting the tables to Excel and then to XML is not unproblematic, however. As the OCR software failed to capture the table content accurately, both steps needed to be done “by hand”, which is error prone. Ideally, XML export should already be possible in pdfs with OCR. Regarding TEI implementation, tables seem to have a shadowy existence so far – or, as TEI pioneer Lou Burnard remarked: “Tables are tricky”. The main reason for this is probably the running text orientation of the existing tools and users. In principle, however, TEI data processing offers the opportunity to think conceptually about the function of data structured in tabular form, and to make changes traceable, especially in serial statistics. This is proved by a project from Basle and Graz using early-modern Basle account books upcycled according to the TEI principles. In addition, the clearly structured text preparation of TEI could provide a training basis to improve the quality of table text recognition.

See extended abstract

 Presented in Session 143. Digitizing Analog History