Interning in the ISAW Library: Extending (and Reimagining) an Error Correction Dataset for Latin Texts
Annotation and dataset building continued in Spring 2026 for the Post-OCR Correction Using Latin Annotations (POCULA) project with three new contributors from our internship program. This semester we approached OCR-correction dataset building with a novel approach, namely we corrected the Latin text of a book of natural history with the additional interest in creating an intermediate student reader of the work. Specifically, we worked on animal descriptions from Conrad Gesner’s Historia Animalium. Interns looked at the corrupted Latin text found in the OCR in scanned editions of Gesner’s work, specifically passages about unicorns, owls, and eels, among other animals. The spring semesters double focus accomplished two goals: 1. anecdotally I can report that this cohort of interns found it more interesting to correct not random error-filled Latin sentences from the project’s target repositories but rather running text about zoological miscellany; and 2. by correcting running text from Gesner, we have also put ourselves in the position of directly using the corrected texts as the basis for an intermediate student reader which I expect to publish in 2027. Win-win, the first a win for Latin NLP tooling, the second for Latin pedagogy.
During the spring semester of 2026, I had the opportunity to contribute to Post-OCR Correction Using Latin Annotations (POCULA) at NYU’s Institute for the Study of the Ancient World. Under Patrick Burns’s guidance, our team helped create training data for a model designed to correct corrupted digital scans of Latin texts. I worked primarily with Conrad Gesner’s Historia Animalium, a sixteenth-century zoological encyclopedia, focusing on animals that interested me, particularly the unicorn (monoceros) and the mole (talpa). My role involved comparing OCR-generated transcriptions against scanned texts and annotating errors so the model could learn from corrections. The project exposed me to Latin far beyond the traditional classical canon of my education and revealed gaps in my understanding of the language’s historical scope.
I also completed annotations of Terence, the 2nd-century BCE comic playwright, a text with which the models struggled even more, particularly with the technical features of dramatic text like formatting dialogue and identifying speakers. Working with so much textual corruption taught me about the intricacies of the scribal tradition and how errors can accumulate through centuries of transmission. Beyond common OCR errors involving letters such as s/f and u/v, I was particularly interested in mistakes caused by early printing conventions, including ligatures such as æ and abbreviations for -que. These features made me more aware of the technical aspects of producing and preserving information in a pre-digital world, which can be easy to overlook in the information age.
Until now, my Latin education has been concentrated on canonical classical texts and traditional methods of study. For me, one of the most valuable aspects of POCULA was seeing how Latin can be integrated into an increasingly digital future. Working toward the publication of a student reader for Gesner was a highlight of my experience. Overall, this project broadened my view of classical reception to include digital humanities.
—Kai Leigh Harrison
During the spring semester, I participated in the Post-OCR Correction Using Latin Annotations (POCULA) project at ISAW with Patrick Burns. The project's goal is to improve the accuracy of OCR software for early printed Latin by creating high-quality annotated training data. My work involved correcting OCR-generated transcriptions, translating passages, and identifying named people (like Aristoteles and Plinius) and places (like Graecia) in Gesner's Historia Animalium.
My first assignment was Gesner's entry on the owl (bubo). At first, I approached the text the obvious way, working from top to bottom and translating every sentence in order. While that was effective for producing accurate annotations, I found it made the reading feel repetitive. Since one goal of projects like POCULA is to make Latin texts more accessible, I realized they should also be engaging to read. When I moved on to Gesner's entry on the parrot (psittacus), I changed my approach. Instead of translating every line sequentially, I selected passages that caught my attention, especially those describing the parrot's ability to imitate human speech, its intelligence, and its unusual behavior. Working this way made the text feel less like an assignment and more like discovering interesting details hidden within a Renaissance encyclopedia.
The project also gave me a new appreciation for the process of preserving historical texts. Many of the corrections involved recognizing early modern printing conventions that OCR software struggled to interpret, including the long s (i.e. ſ), interchangeable u and v, irregular spacing, and archaic spellings. Even small typographical details could completely change the meaning of a sentence. By the end of the semester, I had not only become more comfortable reading Renaissance Latin, but I had also seen how traditional language skills can contribute to modern digital humanities projects by making historical sources more accurate and accessible.
—Kevin O’Neill
Over the spring semester, I worked on the Post-OCR Correction Using Latin Annotations (POCULA) project with Patrick Burns. The project’s goal is to create high-quality training data for a language model capable of correcting (OCR) errors in digitized Latin texts. My work focused on manually correcting corrupted transcriptions and transcribing passages from Conrad Gesner’s Historia Animalium.
Much of my work centered on Gesner’s entries on the eel (anguilis) and the mole (talpa). Unlike the Latin authors I had encountered in school (Caesar, Cicero, Catullus, Horace, etc.), Historia Animalium introduced me to Renaissance natural history. Instead of military campaigns or political speeches, I transcribed descriptions of animal anatomy, the origins of the eel's name, reproduction, methods of capture, culinary practices, and even medical remedies. Translating portions of the text highlighted in particular the contrast between the sixteenth-century understanding of these animals and modern scientific knowledge.
The section on the eel was especially memorable. Gesner discusses everything from the eel's appearance and habitat to competing theories of its reproduction, drawing on authors such as Aristotle. While one passage claims that eels arise spontaneously from decaying matter, another claims that salted eels were considered healthier than fresh ones and warns against eating them if one suffered from arthritis (morbus articularis) or headaches (dolores capitis).
The transcription process itself revealed the challenges of preserving historical texts digitally. I regularly encountered errors involving the long ſ being mistaken for f, u and v being interchanged, missing spaces, and early modern ligatures such as ct, æ, and -que. Over time, these recurring mistakes became recognizable. Additionally, I found it interesting how Gesner’s Latin was interwoven with Greek, Hebrew, French, German, and Italian.
In addition to Historia Animalium, I corrected OCR transcriptions for Act II of Terence’s Andria, where the model struggled even more. It frequently failed to identify speakers correctly, misplaced punctuation, and sometimes produced words that were almost unrecognizable—for example, nunquidnam became *k(unquidnam, reddidisti became *fieddidisti, and nisi became *pjjfi. Comparing these corrupted passages with the original scans demonstrated both the capabilities and limitations of OCR technology. Most importantly, my interventions will become one step toward permanent fixes.
Before this internship, my experience with Latin had been almost entirely confined to the traditional classical canon taught in high school, and it exposed me to the growing field of digital humanities. I am incredibly grateful to Dr. Burns for his guidance throughout the internship. Every discussion—whether about punctuation, speaker abbreviations, or early modern printing—became an opportunity to learn something new. His expertise, encouragement, and sense of humor made my experience both intellectually rewarding and enjoyable. I look forward to seeing how the POCULA project continues to develop!
—Mihou Tatsumi