SCOOP: Source Codes of the Past

2nd Exchange Meeting of SCOOPInternational Network for Automated Text Recognition of Historical Sources

September 7-9, 2026: 2nd exchange meeting, hosted by the Austrian Academy of Sciences (Institute for Medieval Research) and the University of Vienna (Faculty of Historical and Cultural Studies, Institute for Austrian Historical Research)

SCOOP logo

About

SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers; archivists and librarians; computer scientists, machine learning researchers, and software engineers. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well.

Recent years have brought remarkable progress in text recognition technologies, but much work remains in adapting them to the diversity of historical materials, ranging from different scripts and writing systems, varied textual traditions and manuscript structures, to low-resource languages for which training data and tools are scarce. At the same time, ATR is increasingly interwoven with other methodologies, from dataset curation and text reuse analysis to the automation of editorial workflows. These developments are transforming how researchers engage with manuscript sources, and they raise questions that no single discipline can answer alone.

SCOOP exists to bring these conversations together. Technological development, methodological reflection, and the practical needs of editors, cataloguers, and collecting institutions too often happen in separate communities; the network's aim is to let technological and humanistic expertise inform one another directly, across the boundaries of disciplines, languages, and scripts.

How the network works

The work of SCOOP is organised into six working groups, each covering different focus areas. These working groups structure both the meetings and the exchange platform, which supports preparation for the in-person exchange meetings as well as the ongoing discussions and cooperation between the various groups and participants. Each member of SCOOP is a member of one or more working groups.

WG1
HTR Technology Development
WG2
Document or Handwriting Classification
WG3
Methodological Issues of HTR
WG4
Language Challenges
WG5
Datasets and Institutions
WG6
Leveraging Outputs: Text Reuse, NLP and More

2nd Exchange Meeting of SCOOP

After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP exchange meeting will take place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna. Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with an internal working day devoted to the future of SCOOP. Building on these first two meetings, the network is pursuing a durable institutional footing for this collaboration, including a joint European funding application, so that the exchange begun in Princeton can continue to grow.

Monday 7th September

08:00 - 09:00

Registration

PLENARY09:00 - 09:15

Welcome

KEYNOTE09:15 - 10:30

Thibault Clérice ・ ALMAnaCH, Inria

The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition

10:30 - 11:00

Coffee break

Session block 111:00 - 12:30
WG1-1Talks

Technologies & Architectures

  • William Mattingly ・ Yale UniversityFinetuning VLMs in a World of Zero-Shot Frontier Models: A Paradigm Shift toward Knowledge Distillation and Local Deployment
  • Colin Brisson ・ Ecolé pratique des hautes étudesAnandaSky
  • Andy Stauder ・ READ COOPMetatools & Technological Agnosticism
12:30 - 13:30

Lunch

Session block 213:30 - 15:00
Track One
Track Two
Track Three
Track One
WG1-2Talks

Training as a continuous process

  • Doug Emery ・ University of Pennsylvania LibrariesPiloting HTR in the Library: Experiments and Groundwork
  • Wolfgang Göderle ・ University of Graz | University of Passau | MPI GEAMulti-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources
  • Alicia González Martínez ・ Hamburg UniversityFrom PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus
  • Elena Chepel & Anton Repushko ・ University of ViennaAnagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri
Track Two
WG3-1Talks

Transcription approaches I (Palaeography in focus)

  • Benjamin Kiessling ・ ALMAnaCH, Inria ParisComputational Paleography through Automatic Text Recognition
  • Sajjad Nikfahm-Khubravan ・ Roshan Institute for Persian Studies, University of MarylandTowards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts
  • Malamatenia Vlachou ・ IRHT/CNRS-ENPCLeveraging HTR for Script Characterization: Toward a Unified Computational Framework for Palaeographical Analysis
  • Dominique Stutzmann ・ IRHT-CNRS / HU BerlinWorkflows and granularities: text, image, and what remains
Track Three
WG4-1Talks

Language Challenges I

  • Ana Mihaljević ・ Institute for the Croatian languageFour Scripts, Five Languages, One Digitalization Challenge
  • Bernhard Bauer ・ University of GrazMatchbox: Creating a Combined Recognition Model for Early Medieval Celtic Languages and Latin
  • Baptiste Queuche ・ CalfaHow hybrid HTR+VLM approaches are unlocking the processing of under-resourced, non-western languages
  • Seth Kulick ・ Linguistic Data Consortium, University of PennsylvaniaUsing OCR to expand historical treebanks
  • Andrew Janco ・ Princeton UniversityWhich Languages Are In-Vocabulary?
  • Isabelle Marthot-Santaniello ・ University of BaselThe language of Greek papyri? One millennium of writing and its complexity
15:00 - 15:30

Coffee break

Session block 315:30 - 17:00
Track One
Track Two
Track Three
Track One
WG1-3Talks

Experimental approaches

  • Michal Racyn ・ Masaryk UniversityUtilization of the Open WebUI platform in Arkindex
  • Tristan Repolusk ・ Department of Digital Humanities, University of GrazDocument Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26
  • Elise Wang ・ California State University, FullertonFinding our way through the National Archives: HTR for indexing a large corpus
  • Tobias Hodel ・ University of BernInsights from an Unsound Experiment: Testing Kraken, TrOCR, and VLMs with an LLM Judge
Track Two
WG3-2Talks

Transcription approaches II (Transcription decisions, their effect, and related terminology)

  • Jan Odstrčilík ・ Institute for Medieval Research, Austrian Academy of SciencesOtto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR
  • Anna Michalcová ・ Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of SciencesLost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards
  • Ana Mihaljević ・ Institute for the Croatian languageReading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI
  • Paweł Figurski ・ Polish Academy of SciencesHTR of Normalized Latin Texts: Insights from Liturgical Manuscripts
Track Three
WG4-2Talks

Language Challenges II

  • Katrín Lísa van der Linde Mikaelsdóttir ・ University of IcelandDark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic
  • Christine Roughan ・ Princeton UniversityImpacts of Script Features on Text Recognition: Experiments with Greek and Arabic
  • Ephrem Aboud Ishac ・ Institute for Medieval Research, Austrian Academy of SciencesSyriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR
17:00 - 17:15

Break

KEYNOTE17:15 - 18:30

Anna Dolganov & David Smith ・ Austrian Academy of Sciences

New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation (Anna Dolganov and David Smith)

18:30 - 20:00

Reception

Tuesday 8th September

Session block 109:00 - 10:30
Track One
Track Two
Track Three
Track One
WG1-4Roundtable

Roundtable

Track Two
WG2-1Talks

Handwriting Classification I

  • Asimina Paparrigopoulou & Paraskevi Platanou ・ Democritus University of ThraceFrom Handwriting Classification Errors to Paleographic Evidence: Graphic Compensation in Ancient Greek Documentary Hands
  • Tara Andrews ・ University of ViennaClassification of Armenian manuscripts without OCR/HTR preprocessing
  • Serena Ammirati & Paolo Merialdo ・ Università degli Studi Roma TreBeyond the Heatmap: What Faithful Explanations Can Do for Handwriting Identification
Track Three
WG6-1Talks

Leveraging Outputs: Text Reuse, NLP, and More (talks)

  • Alexander O'Neill ・ Musashino UniversityTesting the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts
  • Nikola Krisztian Czindrity ・ University of ViennaFrom Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers
  • Seth Kulick ・ Linguistic Data Consortium, University of PennsylvaniaUsing OCR to expand a treebank of historical Yiddish: A Progress Report
10:30 - 11:00

Coffee Break

Session block 211:00 - 12:30
Track One
Track Two
Track Three
Track One
WG2-2Talks

Handwriting Classification II

  • Aaron Hershkowitz ・ The Institute for Advanced StudyClassifying Squeezes Again: Initial Results from an ICDAR Contest
  • Sebastian Sobecki & Sam Grieggs ・ University of TorontoCommunities of Practice: Capturing the Aspect of Late Medieval Handwriting
  • Giuseppe De Gregorio ・ Computer Vision Center - CVC - BarcelonaScript classification and alphabet identification
Track Two
WG4-3Roundtable

Roundtable 🎤 - Language Challenges

Track Three
WG6-2Demos

Leveraging Outputs: Text Reuse, NLP, and More (demos)

  • Wayne de Fremery ・ Dominican University of CaliforniaNew Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
  • Andrew Janco ・ Princeton UniversityFrom Document Images to Research Catalogue
  • Daniel Tubb ・ Anthropology, University of New Brunswick, CanadaFichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers
  • Martin Roček ・ Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles UniversityIntertextuality as a Retrieval Task: Benchmarking Text Reuse in Classical and Medieval Latin
12:30 - 13:30

Lunch

Session block 313:30 - 15:00
Track One
Track Two
Track Three
Track One
WG3-3Talks

Building ATR/HTR pipelines

  • Michael Schonhardt ・ TU Darmstadt / Akademie der Wissenschaften und der Literatur | MainzThe Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora
  • Olaf Berg ・ Ruhr-Universität BochumBuilding an ATR pipeline for tabular data and script translation
  • Osama Eshera ・ University of MarylandIterative HTR: A new pipeline for automatic text recognition of Arabic-script
  • Johannes Knüchel ・ Austrian National LibraryBuilding an OCR Pipeline at the Austrian National Library
Track Two
WG5-1Talks

Datasets and Institutions

  • Jessie Dummer ・ University of Pennsylvania LibrariesIntegrating HTR into Digital Library Workflows
  • Michael Lužný ・ National Library of the Czech RepublicManuscriptorium Full-Text Module: Building a Digital-Edition Infrastructure
  • Tim Geelhaar ・ Goethe Universität Frankfurt am MainReviving the Legacy: How to Adapt the Latin Text Archive for the Age of AI
  • Ursula Stampfer ・ Bayerische StaatsbibliothekHTR in the German Manuscript Centres: Current Status, Challenges, and Perspectives
Track Three
WG6-3Unconference

Unconference

15:00 - 15:30

Coffee Break

CONCLUSION15:30 - 17:00

Working Groups working on conclusions

BREAK17:00 - 17:15

Short Break

ROUNDTABLE17:15 - 18:45

Final Roundtable - Open data, open code, open minds in AI

CONCLUSION18:45 - 19:00

Conclusion of the public part

Wednesday 9th September

MEETING09:00 - 17:00

Internal meeting of SCOOP

* All times are shown in Central European Summer Time (CEST).

Join the network — and our Discord

SCOOP is open to anyone working on the automatic transcription and analysis of historical sources, across disciplines, scripts, and languages. If you'd like to join the network, write to scoop@oeaw.ac.at, and put Anna Michalcová (a.michalcova@ujc.cas.cz) and Jan Odstrčilík (jan.odstrcilik@oeaw.ac.at) into CC. You will be included into the future communication and given access to the Discord server.

Discord is where much of SCOOP's exchange happens between meetings. It's a free chat platform where conversations are organised into channels: each working group has its own space to discuss its topics, share materials or simply to get to know each other. It's also the easiest place to ask a quick question, share a new tool, dataset, or paper, announce a call or event, or find a collaborator who has already struggled with the same script, language, or pipeline you're facing now.

If you're new to the network, the server is the best way to get a feel for what's going on in SCOOP and to introduce yourself and your project before meeting everyone in person.

Organisers of SCOOP network

  • Tara Andrews

    University of Vienna

    • Scientific committee
  • Anna Dolganov

    Institute for Medieval Research, Austrian Academy of Sciences

    • Scientific committee
  • Gerda Heydemann

    Friedrich-Meinecke-Institut für Geschichte und Historische Kulturwissenschaften

    • Scientific committee
  • Tobias Hodel

    Digital Humanities, University of Bern

    • Scientific committee
  • Alíz Horváth

    Central European University

    • Scientific committee
  • Maria Konstantinidou

    Democritus University of Thrace

    • Scientific committee
  • Anna Michalcová

    Czech Language Institute, Czech Academy of Sciences / Institute for Medieval Research, Austrian Academy of Sciences

    • Local organizing team (Vienna)
    • Scientific committee
  • Jan Odstrčilík

    Institute for Medieval Research, Austrian Academy of Sciences

    • Local organizing team (Vienna)
    • Scientific committee
  • John Pavlopoulos

    Athens University of Economics and Business, and Archimedes, Athena Research Center

    • Scientific committee
  • Paraskevi Platanou

    Athens University of Economics and Business, and Archimedes, Athena Research Center

    • Scientific committee
  • Helmut Reimitz

    Institute for Medieval Research, Austrian Academy of Sciences / Institute for Austrian History Research, University of Vienna

    • Local organizing team (Vienna)
    • Scientific committee
  • Martin Roček

    Institute for Medieval Research, Austrian Academy of Sciences / Faculty of Arts, Charles University

    • Local organizing team (Vienna)
    • Scientific committee
  • Christine Roughan

    Center for Digital Humanities / MARBAS, Princeton University

    • Scientific committee
  • Sofia Torallas Tovar

    Institute for Advanced Study, Princeton

    • Scientific committee
  • Lucia Waldschuetz

    Princeton University / Institute for Medieval Research, Austrian Academy of Sciences

    • Local organizing team (Vienna)
    • Scientific committee
  • Thomas Wallnig

    University of Vienna

    • Local organizing team (Vienna)
    • Scientific committee

Organising institutions