- Python 95.1%
- Shell 4.9%
| .github/workflows | ||
| config | ||
| dist | ||
| scan | ||
| scripts | ||
| tests | ||
| .gitignore | ||
| hints.txt | ||
| ocr_nethack_guide_COMPRESSED.pdf | ||
| README.org | ||
| requirements-scan.txt | ||
intro
This is a scan and OCR of the japanese nethack guide book

download
recommended
-
Primary translated version: Qwen2.5-7B block-level translated hybrid PDF
- This is the best current English facsimile: refined scan backgrounds with Surya OCR blocks replaced by local English translations, validated block-by-block with no missing/overflow rows.
- read the original Japanese OCR online
- Surya Japanese OCR Markdown
- Local Qwen2.5-7B English translation Markdown
review and source artifacts
older or legacy translation artifacts
These are kept for comparison/provenance, but the block-level hybrid PDF above is the recommended English version.
- Older page-level translated facsimile prototype sample PDF
- Side-by-side translation review sample PDF
- legacy uncompressed OCR'ed PDF (google drive link)
- legacy compressed OCR'ed PDF (google drive link)
- legacy compressed OCR'ed PDF (github link)
- legacy DeepL English translation
- legacy Google Translate English translation
Large generated Japanese searchable PDFs and the older full page-level prototype are not committed here because they exceed GitHub's 100MB file limit. Use GitHub Releases for those assembled PDFs; the Build and publish PDF release assets workflow rebuilds them from committed sources and publishes the PDFs as release assets. See dist/README.org for the full generated-artifact layout.
OpenCV scan post-processing pipeline
The non-destructive scan pipeline reads renamed raw scans from scan/, keeps
dist/pdforder.txt as the canonical final page order, and uses
scan/pdforder-rename-manifest.tsv to map final pages back to source images.
The landscape opening scan pdf002-003__nethack_uncropped_openpages_1_2_3-1.png
is split into left/right final pages before deskew/crop. Body pages use the
legacy crop as a raw-scan placement prior, then trim inside the physical page:
extra pixels are removed from the inferred inner/ragged binding edge and from
the scanner-tail-prone bottom edge. Text/rule line segments are preferred for
deskew, because reader-visible baselines matter more than paper-edge angle.
The committed dist/ layout is intentionally organized as:
dist/crops/reference/legacy untrimmed crops used only as stable reference windows.dist/crops/refined/current refined cropped pages for publishing/PDF builds.dist/metadata/refined/metrics and per-page JSON sidecars for the refined crops.dist/text/legacy OCR/translation assets plus local Surya OCR and local Qwen2.5-7B English machine translation outputs.dist/metadata/ocr-bakeoff/OCR engine comparisons and run metadata.dist/metadata/translation-bakeoff/local translation model comparisons.dist/metadata/translation-pdf-prototype/translated facsimile prototype summaries, previews, and Oracle reconstruction guidance.dist/metadata/translation-block-hybrid/block-level translated hybrid PDF source manifest, cleaned-background summary, fit reports, validation notes, and previews.dist/metadata/translation-review-pdf/side-by-side review PDF summaries and preview images.dist/scripts/old one-off shell helpers retained for provenance.
Install the Python/image tooling in an environment of your choice:
python -m pip install -r requirements-scan.txt
Pilot run with debug overlays and the example overrides:
rm -rf work/processed-pilot work/qa-pilot
python scripts/process_scans.py \
--scan-dir scan \
--manifest scan/pdforder-rename-manifest.tsv \
--pdforder dist/pdforder.txt \
--out work/processed-pilot \
--target 1788x2512 \
--reference-crops dist/crops/reference \
--pages 001-020,050,100,150,160,200,240,274 \
--jobs 4 \
--overrides config/scan-process-overrides.example.json \
--debug
python scripts/make_scan_qa.py \
--metrics work/processed-pilot/metrics.jsonl \
--processed-dir work/processed-pilot/color \
--debug-dir work/processed-pilot/debug \
--out work/qa-pilot
python scripts/validate_processed_scans.py \
--pdforder dist/pdforder.txt \
--processed-dir work/processed-pilot/color \
--metrics work/processed-pilot/metrics.jsonl \
--target auto \
--expected-pages 27 \
--pages 001-020,050,100,150,160,200,240,274
Outputs are written under work/processed-pilot/:
color/final color PNG pages, named exactly asdist/pdforder.txtentries.gray_for_ocr/background-normalized grayscale helper PNGs.debug/QA overlays when--debugis used.metadata/per-page JSON sidecars.metrics.jsonlmachine-readable per-page metrics and flags.
QA contact sheets and review.html are written under work/qa-pilot/. Body
pages should all have one output size (currently 1700x2440 with the example
trim defaults); the first cover/front pages may retain their reference sizes. If
a page needs manual correction, add a page or output-name override to a JSON file
based on config/scan-process-overrides.example.json, e.g. angle_deg,
page_side, inner_edge, or trim_insets. For a subset test, write to a
separate output directory such as work/processed-check; --force replaces the
entire output directory and is intentionally restricted to generated
work/processed* directories.
Full run after reviewing pilot output:
rm -rf work/processed work/qa
python scripts/process_scans.py \
--scan-dir scan \
--manifest scan/pdforder-rename-manifest.tsv \
--pdforder dist/pdforder.txt \
--out work/processed \
--target 1788x2512 \
--reference-crops dist/crops/reference \
--jobs 8 \
--overrides config/scan-process-overrides.example.json \
--debug
python scripts/make_scan_qa.py \
--metrics work/processed/metrics.jsonl \
--processed-dir work/processed/color \
--debug-dir work/processed/debug \
--out work/qa
python scripts/validate_processed_scans.py \
--pdforder dist/pdforder.txt \
--processed-dir work/processed/color \
--metrics work/processed/metrics.jsonl \
--target auto \
--expected-pages 274
# Optional stricter structural gate for reviewed severe flags:
python scripts/validate_processed_scans.py \
--pdforder dist/pdforder.txt \
--processed-dir work/processed/color \
--metrics work/processed/metrics.jsonl \
--target auto \
--expected-pages 274 \
--fail-on-flags content_overflow,crop_near_content
The default validator checks order, dimensions, metadata, and metrics. It prints
QA flags but only fails on them when --fail-on-flags is supplied. In refined
trim mode, expected informational flags include reference_crop,
reference_trim, used_line_angle, and sometimes scanner_tail_detected (the
bottom trim is meant to remove those stretched rows from the final crop).
To publish a reviewed full run into dist/crops/refined/:
rsync -a --delete work/processed/color/ dist/crops/refined/
rsync -a --delete work/processed/metadata/ dist/metadata/refined/metadata/
cp work/processed/metrics.jsonl dist/metadata/refined/metrics.jsonl
python scripts/validate_processed_scans.py \
--pdforder dist/pdforder.txt \
--processed-dir dist/crops/refined \
--metrics dist/metadata/refined/metrics.jsonl \
--target auto \
--expected-pages 274
Build a PDF in canonical order. Add --out-ocr-pdf ... --lang jpn+eng only
after installing Japanese Tesseract data (tesseract --list-langs should show
jpn); the build script intentionally does not deskew a second time.
bash scripts/build_processed_pdf.sh \
--processed-dir dist/crops/refined \
--pdforder dist/pdforder.txt \
--out-pdf work/nethack-guide-processed.pdf \
--dpi 300
dist/scripts/croppages.bash is the legacy fixed-crop baseline; the OpenCV
pipeline is preferred for new work. The source scans in scan/ are never
overwritten.