I teach machines to read old and non-European documents, the handwriting and print that standard text recognition tends to neglect.
Most text recognition algorithms and systems are created for modern conventions. Historical material will rarely conform as spelling, layout, and scribal practice vary enormously, and for most collections training data is sparse. My work aims to redesign recognition to cope with that scarcity and messiness: Arabic manuscripts, medieval Hebrew, Chinese inscriptions, really whatever written matter a historian might be interested in.
I also build the infrastructure that makes these methods useful to scholars, chiefly kraken and the eScriptorium platform built on top of it. I try to keep these tools frugal and open. My models run on laptops rather than data centres, shared datasets, and common transcription norms that let scholars build on each other’s work instead of starting from scratch.