- Python 69.8%
- TypeScript 16.3%
- Lua 5.2%
- CSS 4.4%
- TeX 2.7%
- Other 1.5%
| .agents/skills/research-corpus-manifests | ||
| .github/workflows | ||
| .pi | ||
| CME@e1d77c88ca | ||
| config/cme-build | ||
| ConTeXt | ||
| docs | ||
| manifests | ||
| scripts | ||
| tests | ||
| viewer | ||
| .gitignore | ||
| .gitmodules | ||
| AGENTS.md | ||
| by-sa_legaltext.txt | ||
| gpl-3.0.txt | ||
| README.org | ||
- Intro
- Mission
- Technology
- Build and architecture
- Web viewer
- Change-aware validation
- Printing at home
- License
- Resources
Intro
This repository contains Kyvero Press (KVP) production tooling and a working source snapshot drawn from the Corpus of Middle English (CME). Its scope is the source material and publication workflow maintained here; it does not claim to mirror every current CME holding.
Mission
Although most if not all historical texts are in the public domain, great effort and care has gone into transcribing old manuscripts and providing these transcriptions free of charge, and advancements in print on demand technology have made it easier than ever to publish obscure works, it remains nearly impossible to find properly formatted public domain texts at a reasonable price. Instead people are met with either "translations" by large reputable publishing companies or exorbitant prices for unknown quality reprints, many times consisting of low quality scans directly reprinted.
The mission of this project is to provide historical texts in a format that can be printed to reliable aesthetically pleasing and sensible results.
The project aims to provide, first and foremost, a readable edition of the texts in their original languages, accessible for the layperson who is familiar with the language of the text (which, for middle English texts consists of the vast majority of modern English speakers given a brief adjustment period). Later goals could include unofficial critical editions (provided that scholarly parsable public domain critical editions exist), bilingual texts, and interlinear texts.
These resulting texts will be distributed at or near cost via print on demand services in order to ensure the public's access to our shared history. The resulting texts will be freely available and modifiable to all in digital form.
Technology
The current production path normalizes the supported CME XML variants to standalone HTML with Python, passes that HTML through Pandoc, and produces either EPUB or LaTeX typeset as PDF with XeLaTeX. The ConTeXt/ directory contains historical conversion utilities rather than the current production entry point.
Build and architecture
First complete setup and verification. From the repository root, inspect the plan immediately before building one print-profile review candidate under build/:
scripts/cme-build plan CME/source/CME_from_OTA/Gawain.xml \
--profile print-pdf \
--output build/profile/Gawain.pdf
scripts/cme-build single CME/source/CME_from_OTA/Gawain.xml \
--profile print-pdf \
--output build/profile/Gawain.pdf
Artifacts have three stages:
build/: generated candidates, intermediates, and validation evidence;bin/pdf/: a reviewed PDF built with repository defaults for inspection; anddist/: the publication set, replaced only through the fail-closed architecture procedure.
See the build architecture, profile commands, XML formats, and Lulu production checklist for current details.
Web viewer
The React, Vite, and TypeScript viewer under viewer/ turns the metadata and lineage manifests into a searchable catalog, links known source materials, and offers canonical dist/ PDFs for download. Generated viewer data and the deployable static site remain under build/corpus-viewer/.
See the corpus viewer guide for development, testing, catalog generation, and deployment commands.
Change-aware validation
Before pushing manifest or viewer work, run the change-aware gate from the repository root:
python3 scripts/run-changed-gate.py --base origin/main
For ordinary book-manifest changes it validates only the changed work pairs,
their viewer projections, deterministic indexes, and shared vocabulary. It
automatically selects the full corpus gate when shared schemas, validators,
catalog infrastructure, source snapshots, or publication-wide inputs change.
Use --dry-run to inspect the plan or --full to request a periodic complete
audit explicitly.
Printing at home
Home-print imposition and printer submission are unsupported. The repository does not maintain a safe, current printing recipe; use reviewed PDFs with an independently verified local print workflow.
License
All code in this repository is licensed under the GPL-3.0
All output, except that which is in the public domain, is Licensed under CC BY-SA 4.0. For attribution, please provide a link to this repository somewhere in the text, preferably in the colophon, with an explanation that the user may download this and other texts in the corpus.
The reason for licensing under CC BY-SA is to ensure that notice is provided to everyone purchasing the texts that the source and other works are freely available online.
You may use these resulting texts for commercial purposes, and in fact are encouraged to do so.