About
SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers; archivists and librarians; computer scientists, machine learning researchers, and software engineers. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well.
Recent years have brought remarkable progress in text recognition technologies, but much work remains in adapting them to the diversity of historical materials, ranging from different scripts and writing systems, varied textual traditions and manuscript structures, to low-resource languages for which training data and tools are scarce. At the same time, ATR is increasingly interwoven with other methodologies, from dataset curation and text reuse analysis to the automation of editorial workflows. These developments are transforming how researchers engage with manuscript sources, and they raise questions that no single discipline can answer alone.
SCOOP exists to bring these conversations together. Technological development, methodological reflection, and the practical needs of editors, cataloguers, and collecting institutions too often happen in separate communities; the network's aim is to let technological and humanistic expertise inform one another directly, across the boundaries of disciplines, languages, and scripts.
How the network works
The work of SCOOP is organised into six working groups, each covering different focus areas. These working groups structure both the meetings and the exchange platform, which supports preparation for the in-person exchange meetings as well as the ongoing discussions and cooperation between the various groups and participants. Each member of SCOOP is a member of one or more working groups.
- WG1
- HTR Technology Development
- WG2
- Document or Handwriting Classification
- WG3
- Methodological Issues of HTR
- WG4
- Language Challenges
- WG5
- Datasets and Institutions
- WG6
- Leveraging Outputs: Text Reuse, NLP and More
2nd Exchange Meeting of SCOOP
After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP exchange meeting will take place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna. Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with an internal working day devoted to the future of SCOOP. Building on these first two meetings, the network is pursuing a durable institutional footing for this collaboration, including a joint European funding application, so that the exchange begun in Princeton can continue to grow.
Monday 7th September
Registration
Welcome
Thibault Clérice ・ ALMAnaCH, Inria
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition
Coffee break
Technologies & Architectures
- William Mattingly ・ Yale UniversityFinetuning VLMs in a World of Zero-Shot Frontier Models: A Paradigm Shift toward Knowledge Distillation and Local Deployment
- Colin Brisson ・ Ecolé pratique des hautes étudesAnandaSky
- Andy Stauder ・ READ COOPMetatools & Technological Agnosticism
Lunch
Training as a continuous process
- Doug Emery ・ University of Pennsylvania LibrariesPiloting HTR in the Library: Experiments and Groundwork
- Wolfgang Göderle ・ University of Graz | University of Passau | MPI GEAMulti-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources
- Alicia González Martínez ・ Hamburg UniversityFrom PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus
- Elena Chepel & Anton Repushko ・ University of ViennaAnagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri
Transcription approaches I (Palaeography in focus)
- Benjamin Kiessling ・ ALMAnaCH, Inria ParisComputational Paleography through Automatic Text Recognition
- Sajjad Nikfahm-Khubravan ・ Roshan Institute for Persian Studies, University of MarylandTowards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts
- Malamatenia Vlachou ・ IRHT/CNRS-ENPCLeveraging HTR for Script Characterization: Toward a Unified Computational Framework for Palaeographical Analysis
- Dominique Stutzmann ・ IRHT-CNRS / HU BerlinWorkflows and granularities: text, image, and what remains
Language Challenges I
- Ana Mihaljević ・ Institute for the Croatian languageFour Scripts, Five Languages, One Digitalization Challenge
- Bernhard Bauer ・ University of GrazMatchbox: Creating a Combined Recognition Model for Early Medieval Celtic Languages and Latin
- Baptiste Queuche ・ CalfaHow hybrid HTR+VLM approaches are unlocking the processing of under-resourced, non-western languages
- Seth Kulick ・ Linguistic Data Consortium, University of PennsylvaniaUsing OCR to expand historical treebanks
- Andrew Janco ・ Princeton UniversityWhich Languages Are In-Vocabulary?
- Isabelle Marthot-Santaniello ・ University of BaselThe language of Greek papyri? One millennium of writing and its complexity
Coffee break
Experimental approaches
- Michal Racyn ・ Masaryk UniversityUtilization of the Open WebUI platform in Arkindex
- Tristan Repolusk ・ Department of Digital Humanities, University of GrazDocument Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26
- Elise Wang ・ California State University, FullertonFinding our way through the National Archives: HTR for indexing a large corpus
- Tobias Hodel ・ University of BernInsights from an Unsound Experiment: Testing Kraken, TrOCR, and VLMs with an LLM Judge
Transcription approaches II (Transcription decisions, their effect, and related terminology)
- Jan Odstrčilík ・ Institute for Medieval Research, Austrian Academy of SciencesOtto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR
- Anna Michalcová ・ Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of SciencesLost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards
- Ana Mihaljević ・ Institute for the Croatian languageReading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI
- Paweł Figurski ・ Polish Academy of SciencesHTR of Normalized Latin Texts: Insights from Liturgical Manuscripts
Language Challenges II
- Katrín Lísa van der Linde Mikaelsdóttir ・ University of IcelandDark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic
- Christine Roughan ・ Princeton UniversityImpacts of Script Features on Text Recognition: Experiments with Greek and Arabic
- Ephrem Aboud Ishac ・ Institute for Medieval Research, Austrian Academy of SciencesSyriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR
Break
Anna Dolganov & David Smith ・ Austrian Academy of Sciences
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation (Anna Dolganov and David Smith)
Reception
Tuesday 8th September
Roundtable
Handwriting Classification I
- Asimina Paparrigopoulou & Paraskevi Platanou ・ Democritus University of ThraceFrom Handwriting Classification Errors to Paleographic Evidence: Graphic Compensation in Ancient Greek Documentary Hands
- Tara Andrews ・ University of ViennaClassification of Armenian manuscripts without OCR/HTR preprocessing
- Serena Ammirati & Paolo Merialdo ・ Università degli Studi Roma TreBeyond the Heatmap: What Faithful Explanations Can Do for Handwriting Identification
Leveraging Outputs: Text Reuse, NLP, and More (talks)
- Alexander O'Neill ・ Musashino UniversityTesting the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts
- Nikola Krisztian Czindrity ・ University of ViennaFrom Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers
- Seth Kulick ・ Linguistic Data Consortium, University of PennsylvaniaUsing OCR to expand a treebank of historical Yiddish: A Progress Report
Coffee Break
Handwriting Classification II
- Aaron Hershkowitz ・ The Institute for Advanced StudyClassifying Squeezes Again: Initial Results from an ICDAR Contest
- Sebastian Sobecki & Sam Grieggs ・ University of TorontoCommunities of Practice: Capturing the Aspect of Late Medieval Handwriting
- Giuseppe De Gregorio ・ Computer Vision Center - CVC - BarcelonaScript classification and alphabet identification
Roundtable 🎤 - Language Challenges
Leveraging Outputs: Text Reuse, NLP, and More (demos)
- Wayne de Fremery ・ Dominican University of CaliforniaNew Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
- Andrew Janco ・ Princeton UniversityFrom Document Images to Research Catalogue
- Daniel Tubb ・ Anthropology, University of New Brunswick, CanadaFichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers
- Martin Roček ・ Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles UniversityIntertextuality as a Retrieval Task: Benchmarking Text Reuse in Classical and Medieval Latin
Lunch
Building ATR/HTR pipelines
- Michael Schonhardt ・ TU Darmstadt / Akademie der Wissenschaften und der Literatur | MainzThe Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora
- Olaf Berg ・ Ruhr-Universität BochumBuilding an ATR pipeline for tabular data and script translation
- Osama Eshera ・ University of MarylandIterative HTR: A new pipeline for automatic text recognition of Arabic-script
- Johannes Knüchel ・ Austrian National LibraryBuilding an OCR Pipeline at the Austrian National Library
Datasets and Institutions
- Jessie Dummer ・ University of Pennsylvania LibrariesIntegrating HTR into Digital Library Workflows
- Michael Lužný ・ National Library of the Czech RepublicManuscriptorium Full-Text Module: Building a Digital-Edition Infrastructure
- Tim Geelhaar ・ Goethe Universität Frankfurt am MainReviving the Legacy: How to Adapt the Latin Text Archive for the Age of AI
- Ursula Stampfer ・ Bayerische StaatsbibliothekHTR in the German Manuscript Centres: Current Status, Challenges, and Perspectives
Unconference
Coffee Break
Working Groups working on conclusions
Short Break
Final Roundtable - Open data, open code, open minds in AI
Conclusion of the public part
Wednesday 9th September
Internal meeting of SCOOP
* All times are shown in Central European Summer Time (CEST).
Join the network — and our Discord
SCOOP is open to anyone working on the automatic transcription and analysis of historical sources, across disciplines, scripts, and languages. If you'd like to join the network, write to scoop@oeaw.ac.at, and put Anna Michalcová (a.michalcova@ujc.cas.cz) and Jan Odstrčilík (jan.odstrcilik@oeaw.ac.at) into CC. You will be included into the future communication and given access to the Discord server.
Discord is where much of SCOOP's exchange happens between meetings. It's a free chat platform where conversations are organised into channels: each working group has its own space to discuss its topics, share materials or simply to get to know each other. It's also the easiest place to ask a quick question, share a new tool, dataset, or paper, announce a call or event, or find a collaborator who has already struggled with the same script, language, or pipeline you're facing now.
If you're new to the network, the server is the best way to get a feel for what's going on in SCOOP and to introduce yourself and your project before meeting everyone in person.
Organisers of SCOOP network
Tara Andrews
University of Vienna
- Scientific committee
Anna Dolganov
Institute for Medieval Research, Austrian Academy of Sciences
- Scientific committee
Gerda Heydemann
Friedrich-Meinecke-Institut für Geschichte und Historische Kulturwissenschaften
- Scientific committee
Tobias Hodel
Digital Humanities, University of Bern
- Scientific committee
Alíz Horváth
Central European University
- Scientific committee
Maria Konstantinidou
Democritus University of Thrace
- Scientific committee
Anna Michalcová
Czech Language Institute, Czech Academy of Sciences / Institute for Medieval Research, Austrian Academy of Sciences
- Local organizing team (Vienna)
- Scientific committee
Jan Odstrčilík
Institute for Medieval Research, Austrian Academy of Sciences
- Local organizing team (Vienna)
- Scientific committee
John Pavlopoulos
Athens University of Economics and Business, and Archimedes, Athena Research Center
- Scientific committee
Paraskevi Platanou
Athens University of Economics and Business, and Archimedes, Athena Research Center
- Scientific committee
Helmut Reimitz
Institute for Medieval Research, Austrian Academy of Sciences / Institute for Austrian History Research, University of Vienna
- Local organizing team (Vienna)
- Scientific committee
Martin Roček
Institute for Medieval Research, Austrian Academy of Sciences / Faculty of Arts, Charles University
- Local organizing team (Vienna)
- Scientific committee
Christine Roughan
Center for Digital Humanities / MARBAS, Princeton University
- Scientific committee
Sofia Torallas Tovar
Institute for Advanced Study, Princeton
- Scientific committee
Lucia Waldschuetz
Princeton University / Institute for Medieval Research, Austrian Academy of Sciences
- Local organizing team (Vienna)
- Scientific committee
Thomas Wallnig
University of Vienna
- Local organizing team (Vienna)
- Scientific committee
Organising institutions
Institute for Medieval Research
www.oeaw.ac.at/en/imafo/homeAustrian Academy of Sciences
Institute for Austrian Historical Research
geschichtsforschung.univie.ac.atUniversity of Vienna
Faculty of Historical and Cultural Study
hist-kult.univie.ac.atUniversity of Vienna
Department of History
ifg.univie.ac.at/enUniversity of Vienna
Center for Digital Humanities
cdh.princeton.eduPrinceton University
Humanities Initiative
initiative.humanities.princeton.eduPrinceton University
School of Historical Studies
www.ias.edu/hsInstitute for Advanced Study, Princeton





