coOCR/HTR
Recognises the text of early printed books and manuscripts and lets every reading be checked against the page image.
With frontier language models and coding agents we develop tools, workflows and knowledge bases for your work processes, on commission, together with your team or in training. The basis is a precise understanding of your project and your data.
In workshops and intensive days your team acquires the competence to work with coding agents and knowledge bases itself.
We develop together with your team on your data. The knowledge base remains with you and enables independent further development.
We develop the tool on your behalf, from a small application for a single work step to an agentic pipeline.
Frontier language models are used in two places, as coding agents during development and in the finished tool for individual work steps such as text recognition, markup or the generation of structured data.
In both cases commercial as well as open models come into consideration, up to operation on your own hardware. Model and access are chosen together with you according to the nature and protection needs of the data. For development a few sample files, anonymised where required, are usually sufficient.
Billing depends on the task, by the hour, by working day or as a flat fee. The starting point is your budget. After an initial assessment we set out what is possible within it, and you decide after each step whether to continue.
Recognises the text of early printed books and manuscripts and lets every reading be checked against the page image.
Edits TEI XML as readable text and validates against the schema on saving.
Shows letter metadata from CMIF as a map, a timeline and a network, so that it becomes visible who wrote to whom, when and from where.
Transcribes an entire estate and records for every object whether a machine, an agent or a human has checked it.
Rescues a bibliography from a decommissioned wiki into research data in which every statement keeps its source page.
Promptotyping develops research artefacts with AI agents out of a maintained knowledge base.
Grounded Vault leads every load-bearing claim in a report through a verifiable anchor back to the source passage.
The Agentic Edition Pipeline takes digital copies through transcription and review to TEI and a static reading view.
Research Mission Control divides clarifying, implementing and checking among AI agents working on a shared repository.
Work begins with an understanding of the project, its research question, its data and its workflows. This knowledge is recorded in a project knowledge base and versioned with the code. The method of context engineering follows from it, for example Promptotyping, Grounded Vault, an agentic pipeline or a simple work cycle with one agent.
The quality of agent-generated code depends on the context in which the agents work. We design this framework from the knowledge base, precise requirements, sample data from your holdings and automated tests that the agents run themselves.
Results are prototypes and research tools whose maturity is openly declared. Productive operation, for example with sensitive data or many users, requires professional revision and an independent review of the code.
Results are held in open, documented formats, depending on the material TEI, PAGE XML, METS/MODS, JSON-LD or CSV. Image data can be integrated via IIIF. The data thus remain readable independently of the tool and can be transferred to other systems and repositories.
Where large language models are used, it remains discernible for every item whether it was generated by a model, checked by an agent or confirmed by your experts. Only content confirmed by your experts counts as established.
You receive the complete source code and the project knowledge base, on the basis of which your team, external developers or AI assistants can continue the work. For research data we offer long-term archiving in the certified repository GAMS.
Digital Humanities Craft is a company that grew out of research in the Digital Humanities, based near Graz. We develop research software and digital editions, teach at universities in Austria and Germany and work as a technical partner in funded projects. Team and projects
Using real files we discuss your workflow and clarify what is possible within the intended budget.
Goal, data and acceptance criteria are recorded in writing.
A first working version is tried out in daily work.
Your experts assess the result against the agreed criteria.
Tool, source code and knowledge base are handed over.
Please describe your workflow briefly. The following details are helpful:
office@dhcraft.org
Send us what you already have. From it we build a first prototype on which the project can be discussed in concrete terms.
office@dhcraft.org