Automated sign detection across the Electronic Babylonian Library
The authors present a large‑scale cuneiform sign detection system that uses a DETR‑based model trained on the biggest annotated dataset to date. The pipeline combines automatic tablet extraction, line grouping, and n‑gram similarity, achieving 28‑37% improvements over previous COCO‑style metrics. It was applied to over 87,000 tablet fragments from the Electronic Babylonian Library, generating nearly 2.9 million sign detections and offering a scalable foundation for future multimodal analysis.
- ▪The study introduces a DETR‑based object detection model evaluated with 173 and 106 sign classes.
- ▪The system integrates tablet‑side extraction, heuristic line grouping, and n‑gram textual similarity to link visual detection with textual structure.
- ▪Performance gains of 28‑37% over prior work are reported on COCO‑style detection metrics.
- ▪The method processed 87,668 tablet fragments, producing approximately 2.9 million sign detections without relying on linguistic priors.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Computer Vision and Pattern Recognition arXiv:2606.22608 (cs) [Submitted on 21 Jun 2026] Title:Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline Authors:Wentao Che, Esteban Garcés Arias, Asim Niaz, Andreas Bender, Enrique Jiménez View a PDF of the paper titled Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline, by Wentao Che and 4 other authors View PDF HTML (experimental) Abstract:Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.