u/HauntingInstance9

Handwriting recognition on 19th century manuscript material - has anyone benchmarked this against Transkribus?

We've been working on handwritten text recognition for historical material and I'm interested in how people here are handling transcription backlogs.

The specific problem- collections that were never going to be transcribed by hand at any realistic staffing level. A scanned document that can't be read is functionally invisible it's in the catalogue, it's preserved, and nobody can search it.

Attached is a run on an 1856 US land record: transcription, then translation, then export to plain text. Roughly 30 seconds end to end.

What I'd like to know from people doing this at scale:

  1. What accuracy threshold do you consider acceptable before machine transcription goes into a public catalogue?

  2. Are you publishing machine output flagged as unverified, or holding it until a human checks it?

  3. Where does it break for you - secretary hand, Kurrent, marginalia, or something else?

Disclosure: I work with Scripily, one of the tools in this space. Genuinely asking about practice here.

reddit.com
u/HauntingInstance9 — 8 days ago