Handwriting recognition on 19th century manuscript material - has anyone benchmarked this against Transkribus?
We've been working on handwritten text recognition for historical material and I'm interested in how people here are handling transcription backlogs.
The specific problem- collections that were never going to be transcribed by hand at any realistic staffing level. A scanned document that can't be read is functionally invisible it's in the catalogue, it's preserved, and nobody can search it.
Attached is a run on an 1856 US land record: transcription, then translation, then export to plain text. Roughly 30 seconds end to end.
What I'd like to know from people doing this at scale:
What accuracy threshold do you consider acceptable before machine transcription goes into a public catalogue?
Are you publishing machine output flagged as unverified, or holding it until a human checks it?
Where does it break for you - secretary hand, Kurrent, marginalia, or something else?
Disclosure: I work with Scripily, one of the tools in this space. Genuinely asking about practice here.