25 September 2026

Knowledge, Context and Agentic Engineering for Research Data Workflows & Digital Editions

Machine Learning for Digital Scholarly Editions

Event CLARIAH-AT Summer School 2026 Language English Licence CC BY 4.0

Abstract

This one-day workshop examines how language models and AI harnesses can support source-bound research data workflows and digital scholarly editions. A controlled sequence of model calls shows how research questions, segment boundaries, schemas and evidence requirements change an extraction result. Two hands-on units then move from multimodal transcription of a manuscript page to structured information extraction, with explicit validation and scholarly status.

Learning objectives

  • Distinguish prompt, knowledge, context and agentic engineering within a research workflow.
  • Explain why segment boundaries and schemas are research decisions.
  • Assess a generated transcription against its source.
  • Record machine output, human validation and scholarly acceptance as separate states.

Outline

  1. Research data workflows, language models and AI harnesses
  2. Four model calls with progressively explicit context and constraints
  3. Multimodal transcription of a manuscript page
  4. Schema-based extraction and scholarly verification

Download Slide Deck as PDF

Literature and references

← All workshops