Corpus of ancient and medieval texts https://kyvero-press.github.io/General-Corpus/
  • Python 69.8%
  • TypeScript 16.3%
  • Lua 5.2%
  • CSS 4.4%
  • TeX 2.7%
  • Other 1.5%
Find a file
2026-07-22 00:09:53 -04:00
.agents/skills/research-corpus-manifests Add rollewks English prose manifests 2026-07-18 01:09:22 -04:00
.github/workflows Publish Pages from deployment branch 2026-07-19 20:31:17 -04:00
.pi Add General Corpus PDF task graph 2026-06-16 21:39:33 -04:00
CME@e1d77c88ca add CME source submodule 2024-12-22 02:27:20 -05:00
config/cme-build Add Lulu production PDF support 2026-07-07 17:35:34 -04:00
ConTeXt parse titles into subtitles using lua, switch to unbalanced columns, use dropcaps at beginning of paragraphs, add script to auto recompile, penalize widow and orphan creation 2024-12-29 08:09:57 -05:00
docs Allow cacheless Pages manifest validation 2026-07-22 00:09:53 -04:00
manifests Add rollewks English prose manifests 2026-07-18 01:09:22 -04:00
scripts Allow cacheless Pages manifest validation 2026-07-22 00:09:53 -04:00
tests Allow cacheless Pages manifest validation 2026-07-22 00:09:53 -04:00
viewer Allow cacheless Pages manifest validation 2026-07-22 00:09:53 -04:00
.gitignore Track local source cache availability 2026-07-11 18:15:13 -04:00
.gitmodules add CME source submodule 2024-12-22 02:27:20 -05:00
AGENTS.md Add researched corpus metadata and lineage 2026-07-11 16:32:01 -04:00
by-sa_legaltext.txt add licenses 2024-12-22 02:36:59 -05:00
gpl-3.0.txt add licenses 2024-12-22 02:36:59 -05:00
README.org Add change-aware validation gate 2026-07-12 18:07:35 -04:00

Intro

This repository contains Kyvero Press (KVP) production tooling and a working source snapshot drawn from the Corpus of Middle English (CME). Its scope is the source material and publication workflow maintained here; it does not claim to mirror every current CME holding.

Mission

Although most if not all historical texts are in the public domain, great effort and care has gone into transcribing old manuscripts and providing these transcriptions free of charge, and advancements in print on demand technology have made it easier than ever to publish obscure works, it remains nearly impossible to find properly formatted public domain texts at a reasonable price. Instead people are met with either "translations" by large reputable publishing companies or exorbitant prices for unknown quality reprints, many times consisting of low quality scans directly reprinted.

The mission of this project is to provide historical texts in a format that can be printed to reliable aesthetically pleasing and sensible results.

The project aims to provide, first and foremost, a readable edition of the texts in their original languages, accessible for the layperson who is familiar with the language of the text (which, for middle English texts consists of the vast majority of modern English speakers given a brief adjustment period). Later goals could include unofficial critical editions (provided that scholarly parsable public domain critical editions exist), bilingual texts, and interlinear texts.

These resulting texts will be distributed at or near cost via print on demand services in order to ensure the public's access to our shared history. The resulting texts will be freely available and modifiable to all in digital form.

Technology

The current production path normalizes the supported CME XML variants to standalone HTML with Python, passes that HTML through Pandoc, and produces either EPUB or LaTeX typeset as PDF with XeLaTeX. The ConTeXt/ directory contains historical conversion utilities rather than the current production entry point.

Build and architecture

First complete setup and verification. From the repository root, inspect the plan immediately before building one print-profile review candidate under build/:

  scripts/cme-build plan CME/source/CME_from_OTA/Gawain.xml \
    --profile print-pdf \
    --output build/profile/Gawain.pdf
  scripts/cme-build single CME/source/CME_from_OTA/Gawain.xml \
    --profile print-pdf \
    --output build/profile/Gawain.pdf

Artifacts have three stages:

  • build/: generated candidates, intermediates, and validation evidence;
  • bin/pdf/: a reviewed PDF built with repository defaults for inspection; and
  • dist/: the publication set, replaced only through the fail-closed architecture procedure.

See the build architecture, profile commands, XML formats, and Lulu production checklist for current details.

Web viewer

The React, Vite, and TypeScript viewer under viewer/ turns the metadata and lineage manifests into a searchable catalog, links known source materials, and offers canonical dist/ PDFs for download. Generated viewer data and the deployable static site remain under build/corpus-viewer/.

See the corpus viewer guide for development, testing, catalog generation, and deployment commands.

Change-aware validation

Before pushing manifest or viewer work, run the change-aware gate from the repository root:

  python3 scripts/run-changed-gate.py --base origin/main

For ordinary book-manifest changes it validates only the changed work pairs, their viewer projections, deterministic indexes, and shared vocabulary. It automatically selects the full corpus gate when shared schemas, validators, catalog infrastructure, source snapshots, or publication-wide inputs change. Use --dry-run to inspect the plan or --full to request a periodic complete audit explicitly.

Printing at home

Home-print imposition and printer submission are unsupported. The repository does not maintain a safe, current printing recipe; use reviewed PDFs with an independently verified local print workflow.

License

All code in this repository is licensed under the GPL-3.0

All output, except that which is in the public domain, is Licensed under CC BY-SA 4.0. For attribution, please provide a link to this repository somewhere in the text, preferably in the colophon, with an explanation that the user may download this and other texts in the corpus.

The reason for licensing under CC BY-SA is to ensure that notice is provided to everyone purchasing the texts that the source and other works are freely available online.

You may use these resulting texts for commercial purposes, and in fact are encouraged to do so.