16 and 17 September 2026

Making Estate Materials Digitally Accessible

Research Data Workflows with Large Language Models and Promptotyping

Event KUG Summer School 2026, University of Music and Performing Arts Graz Language English Licence CC BY 4.0

Abstract

Across four connected sessions, the workshop follows estate material from digitisation to a usable digital research artefact. The Wednesday session uses Stefan Zweig Digital to explain metadata, TEI XML, data modelling, authority data, repository ingest and publication. On Thursday, participants work with M³GIM objects, examine the role of large language models, extract and validate structured data, and develop a small browser-based application through Promptotyping.

Learning objectives

  • Explain the transformations between a digitised source and a published research resource.
  • Identify suitable roles for large language models and define the required checks.
  • Create structured data while recording source relations, provenance and uncertainty.
  • Translate a reduced research question and checked data into application requirements.

Outline

  1. The Stefan Zweig Digital workflow from archival object to research platform
  2. Large language models in research data workflows
  3. OCR and schema-based information extraction with M³GIM objects
  4. Promptotyping a digital research application

Web Module

This browser-native sequence follows one estate object through five accountable transformations. Use the unit links to move directly to a topic; each unit ends with a result that can be inspected.

Unit 01

From estate object to research platform

A digitised image remains one representation of an archival object. Description, transcription, entity modelling, repository ingest and publication create further representations with distinct functions. Stable identifiers and provenance connect them while preserving their differences.

The workflow becomes accountable when every transition names its input, transformation, output and check. This also reveals which questions an interface can answer: a map requires identified places and explicit relations, while a timeline additionally requires normalised dates and visible uncertainty.

Observable check: Select one platform statement and trace it back through the repository record to the identified source object.

Unit 02

Where language models can assist

A language model can propose transcriptions, structured values, authority matches or format conversions. Its fluency supplies no evidence that a passage is complete, a relation is supported or an identifier denotes the intended entity. Model output therefore remains a candidate transformation.

AI readiness belongs to the workflow. A usable task needs permitted and legible input, an explicit output contract, representative examples and a verification route. Rights, confidentiality, prompt records, model settings and corrections remain part of the documented process.

Observable check: For each proposed model task, name the source or rule against which its output will be checked.

Unit 03

From source to checked structured data

Transcription and information extraction are separate transitions. A diplomatic transcription first records what is visible, including line structure and uncertain readings. Structured extraction then selects entities, properties and relations according to a documented schema.

The data record keeps the object identifier, the supporting passage and its review status. Missing values remain empty, and interpretations occupy a separate labelled field. Candidate authority matches are compared with project data and the consulted authority record before an identifier is accepted.

Observable check: Every populated value in the synthetic record points to a source passage and carries a review status.

Unit 04

From research question to browser application

Promptotyping begins with a bounded research question and a compact knowledge base. The knowledge base defines the available data, entity relations, rights, uncertainty rules, intended users and acceptance criteria. These maintained statements guide implementation more reliably than a single long prompt.

The first browser artefact uses the smallest dataset and feature set that can test the research operation. A movement view, for example, needs identified places, typed relations and temporal information. The interface must retain links to source evidence and avoid precision that the data cannot support.

Observable check: Each view and filter names the data fields it uses, and each displayed claim can be traced to an object record.

Unit 05

Verification and responsible acceptance

Validation, verification and acceptance answer different questions. Automated checks can confirm syntax, schema conformance, identifiers and package structure. Scholarly verification compares a representation with its source, evaluates interpretation and records uncertainty.

A responsible person then decides whether the result is adequate for the defined research purpose and assigns its status. A technically valid file may fail source-based review, while a verified statement may remain restricted by rights or repository policy.

Observable check: The validation record identifies the performed tests, the reviewed evidence, unresolved limits and the role responsible for acceptance.

Download Slide Deck as PDF

Literature and references

← All workshops