This is the born-digital stage of the ATRIUM pipeline. A document that was created on a
computer already holds its text. digital-convert reads that text, with its page structure,
reading order, headings, tables and (for PDF) exact geometry, and writes it as an
atrium_document record. The record is the JSON every other ATRIUM stage accretes onto.
No OCR, no model, no network.
- Inputs, recognised by content: PDF, DOCX, ODT, ODS, XLSX, RTF, and the legacy DOC and XLS through headless LibreOffice.
- Output: the record (
source,pages,lines,content,tables,provenance), validated against the frozen schema. On request it is also rendered as annotated Markdown. - Two ways to run it: the command line (
api_util/digital_to_json.py) and the HTTP serviceapi-digital(service/api.py), which has two endpoints:POST /reformat: file in, record out. This is the endpoint the AMΔR route calls. It calls no other service; in the AMΔR deployment Temporal chains the stages.POST /describe: the same conversion, plus a per-page assessment. For every page it reports:- whether the embedded text is usable;
- if it is not, what kind of page it is (page-classification);
- whether the text reads (ocr-postprocess);
- where the page goes next (NLP, OCR or HTR);
- its layout and its text. It calls the two stages only when their URLs are configured; until they are deployed, it reports them as not configured and answers from the converter alone.
Note
This repository was atrium-llm-enrich until v1.0.0-beta.
- The keyword stage it used to hold (the LLM over the AMΔR/TEATER vocabularies,
/extract_keywords) moved to atrium-keyword-extract, see #1. The removed files can be recovered from tagv1.0.0-beta; the transfer manifest is inagent_dev_logs/digests/1.digest.md. - Since v1.1.0-beta the service id is
atrium-digital-convert(the spec declares the rename,x-atrium-service-previous: atrium-llm-enrich), the images areghcr.io/ufal/atrium-digital-convert-{api,digital}, and the record program id staysdigital-convert.
- Where it sits in the pipeline
- βοΈ Setup
- Command line (
api_util/digital_to_json.py) - Formats
- What the record holds
- The service:
/reformatand/describe - The per-page assessment
- Record β annotated Markdown
- Configuration
- π³ Docker
- Paradata and provenance
- π Document Understanding benchmark (research tools)
- Development
- Acknowledgements
AMΔR upload βββΊ seed record (doc_id, source.sha512, filename, media_type)
β
βΌ
born-digital? ββyesβββΊ digital-convert /reformat βββΊ record βββΊ nlp-enrich, keyword-extract, β¦
β β
no ββ pages flagged needs_ocr βββΊ the OCR/HTR route
βΌ (ocr-postprocess merges the re-OCR'd pages)
OCR / HTR route
- The record's originator. For a
digital-born-*record, digital-convert alone writes the positional plane (pages,lines,content,tables).atrium_document.ORIGIN_ORIGINATORSenforces that rule. Every later stage adds its own blocks to the same record. - The AMΔR seed. With a seed (
document_json), the record keeps the seed'sdoc_idand source facts. The seed'ssource.sha512is checked against the uploaded bytes before anything is parsed: a mismatch issource_digest_mismatch(HTTP 422, CLI exit 3), so a record is never attached to the wrong file. - Pages the converter cannot vouch for are flagged
pages[].needs_ocr, with a reason, and every page says why inpages[].text_layer(digital,garbled,ocr,none,blank):- no text layer (a scan, text drawn as curves):
none; - a text layer that does not decode (broken font encodings, replacement characters):
garbled; - a few pages that are a prior OCR run:
ocr. The record is still written, and those pages go to the OCR/HTR route. atrium-ocr-postprocess merges the ATR ALTO of each such page back into the same record, that page only. A PDF that is mostly a prior OCR run (at leastOCR_LAYER_DOCUMENT_SHAREof its pages, default 0.5) is not born-digital. It is refused asocr_text_layerand belongs to the OCR route as a whole.
- no text layer (a scan, text drawn as curves):
- The legacy DOC and XLS were accepted by AMΔR on #4 (2026-09-26). LibreOffice converts the file to DOCX/XLSX first and is not counted in the record's licence; see Paradata and provenance.
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt # base: pydantic, requests, jsonschema, lxml, β¦
pip install -r requirements_digital.txt # the converter: pdfplumber, pypdfium2, python-docx, lxml, jsonschema
pip install -r service/requirements.txt # + the web server (FastAPI, uvicorn), for the serviceOptional:
pip install -r requirements_digital_docling.txt # the heavy PDF engine (--engine docling): Docling + torch
apt-get install libreoffice-writer-nogui libreoffice-calc-nogui # DOC and XLS (or set LIBREOFFICE_BIN)
pip install -r requirements_flexiconv.txt # flexiconv adapter (GPL-3.0; api_util/flexiconv_convert.py)
pip install -r requirements_docmd.txt # deprecated direct converters (--legacy, --ocr)python3 api_util/digital_to_json.py report.pdf --document-json-out report.document.json
python3 api_util/digital_to_json.py report.odt --document-json seed.json --document-json-out report.document.json
python3 api_util/digital_to_json.py budget.xls --out-dir records/
python3 api_util/digital_to_json.py report.pdf --engine docling --paradata-dir paradata/| Option | Meaning |
|---|---|
--document-json-out (--out) |
where to write the record (default <out-dir>/<doc_id>.document.json) |
--document-json |
the AMΔR seed, or an earlier record of the same document, to accrete onto |
--doc-id |
override the doc_id derived from the file name (a seed's doc_id always wins) |
--engine |
light (default) or docling (PDF only, needs requirements_digital_docling.txt and its models) |
--docx-page-breaks |
DOCX pages: auto (explicit, section and Word's last rendered breaks), explicit, none |
--paradata-dir |
also write the run's paradata record |
--strict |
raise instead of warning on a field-ownership violation |
Exit codes:
0: the record was written.2: a dependency is missing (a reader, Docling's models, LibreOffice for DOC/XLS), with advice.3: not an input this converter takes (unsupported type, an OCR-layer PDF, an old binary Office type it cannot convert), or a seed whosesource.sha512does not match.4: corrupt, encrypted, over the ZIP limits, a failed LibreOffice conversion, or over a limit (MAX_PAGES,LIBREOFFICE_TIMEOUT_S).
The type is decided by the file's content. The extension only names the file, so a mislabelled upload is read as what it is.
| Input | Recognised by | Reader | source.origin |
Pages | Geometry |
|---|---|---|---|---|---|
%PDF- |
digital_pdf.py (pdfplumber, pypdfium2), or Docling (--engine docling) |
digital-born-pdf |
PDF pages, named by /PageLabels (i, ii, A-1) |
exact boxes, points, top-left origin | |
| DOCX | ZIP with word/document.xml |
digital_docx.py (python-docx, lxml) |
digital-born-docx |
explicit, section and Word's last rendered page breaks | none |
| ODT | ZIP, ODF mimetype |
text_formats.py, the shared reader (vendored from atrium-ocr-postprocess) |
digital-born-odt |
explicit and soft page breaks | none |
| ODS | ZIP, ODF mimetype |
text_formats.py |
digital-born-ods |
one per sheet, labelled by the sheet name | none |
| XLSX | ZIP with xl/workbook.xml |
text_formats.py |
digital-born-xlsx |
one per sheet, labelled by the sheet name | none |
| RTF | {\rtf |
text_formats.py |
digital-born-rtf |
\page breaks |
none |
| DOC | OLE2 with a WordDocument stream |
headless LibreOffice β DOCX β digital_docx.py |
digital-born-doc |
as DOCX | none |
| XLS | OLE2 with a Workbook/Book stream |
headless LibreOffice β XLSX β text_formats.py |
digital-born-xls |
one per sheet | none |
- Refused with a reason:
- PPTX, ODP, EPUB, plain text and any other type:
unsupported, HTTP 415, which lists the accepted types; - an encrypted file:
encrypted; - a broken one:
corrupt; - a ZIP bomb:
zip_limits_exceeded.
- PPTX, ODP, EPUB, plain text and any other type:
- Only PDF has geometry. The other formats have no page coordinates without rendering them, so
their records carry no
bboxand nocanvas. That follows the schema's rule against fabricated boxes. - Spreadsheets. A row's text cells become one line, joined by a tab. Numbers, dates and
formulas are not text and are dropped by the reader. Each sheet is one
group_id, so a consumer that builds paragraphs from groups keeps a sheet together. Spreadsheettables[](cell grids) are a follow-up. - The shared reader. ODT, ODS, XLSX and RTF go through atrium-ocr-postprocess's reader (the hub's
"one document reader", atrium-project#72), not a second fork.
tests/test_vendored_reader_parity.pypins the copy by its SHA-256. The para-drift workflow compares it with ocr-postprocess'stesthead, and the hub'sscripts/revendor_shared.shrefreshes it. - DOC/XLS conversion:
- LibreOffice runs headless, with a private, throw-away profile and a time limit
(
LIBREOFFICE_TIMEOUT_S).LIBREOFFICE_BINnames the binary (default:sofficeorlibreofficeonPATH). - The record keeps the original file's name, media type and digest, so it describes the DOC/XLS that was uploaded, not the converted file.
- Without LibreOffice a DOC/XLS fails with
dependency_missing(HTTP 501, CLI exit 2); every other format still works.
- LibreOffice runs headless, with a private, throw-away profile and a time limit
(
| Layout cue | PDF (light engine) | DOCX | ODT / RTF | ODS / XLSX | Record field |
|---|---|---|---|---|---|
| Pages | PDF pages, named by /PageLabels |
explicit, section and Word's last rendered page breaks | page breaks | sheets | pages[] (page, page_index) |
| Text blocks / paragraphs | vertical gaps, per column | one per paragraph, one per table cell | paragraph, table cell | row | lines[], lines[].group_id |
| Reading order | words, columns (column-major), headers first | document order; text boxes after their anchor | document order | row order | lines[] order |
| Headings | font size against the body size | outline level, then style name (Heading N, Nadpis N) |
β | β | lines[].style.heading_level |
| Bold / italic | font name | run β character style β paragraph style | β | β | lines[].style.bold / .italic |
| Running header / footer | margin lines repeated across pages, page numbers | each section's header and footer | β | β | lines[].style.region |
| Footnotes | β (Docling engine: yes) | footnotes and endnotes | β | β | lines[].style.region = footnote |
| Tables | ruled tables (Docling engine: any) | tables, merged cells | cells as lines | β (follow-up) | tables[] + cells[].group_id |
| Bounding boxes, page size | exact, points, top-left origin | none | none | none | lines[].bbox, pages[].canvas |
| Untrustworthy text | mojibake, replacement characters, no text layer, prior OCR | the same decode check | the same decode check | the same decode check | lines[].categ = Garbage, pages[].needs_ocr |
- Engines.
--engine light(default) usesrequirements_digital.txt: permissive licences, no models, no network.--engine doclingis opt-in and PDF only (requirements_digital_docling.txt). Docling's layout and table models decide reading order, headings, furniture, footnotes and tables on complex pages, and the light engine's lines keep their exact geometry.- Docling needs its model weights:
docling-tools models download layout tableformer -o DIRandDOCLING_ARTIFACTS_PATH=DIR. The Docker targetdigital-doclingdoes this at build time; it is built locally, never published. OCR stays off in both engines.
- Output gate. A record is never written if it fails any of these:
- the field-ownership round trip (only digital-convert's own fields);
- the JSON Schema;
- the seed digest check.
pip install -r service/requirements.txt
python -m service.api # 0.0.0.0:8000
curl -F file=@report.pdf http://localhost:8000/reformat
curl -F file=@report.pdf -F document_json=@seed.json -F markdown=true http://localhost:8000/reformat
curl -F file=@report.pdf http://localhost:8000/describe| Endpoint | Does | Calls other services |
|---|---|---|
POST /reformat |
file (+ seed) β document_json (the record), Markdown on request, paradata |
never |
POST /describe |
the same conversion β document_json, pages[] (the assessment), summary, stages[], paradata |
page-classification and ocr-postprocess, only when configured |
GET /info, /health, /ready |
the ATRIUM service contract: identity, limits, readers, LibreOffice availability, stages configured | β |
The fields, every response key, the error table (ocr_text_layer, source_digest_mismatch,
unsupported_media_type, limit_exceeded, busy, β¦), the limits and the shutdown behaviour are
documented in service/README.md π. The typed contract is
service/openapi.json π, attached to every release.
/describe answers, page by page:
- can the embedded text be used as it is?
- if not, what kind of page is it?
- does the text read?
- where should the page go next?
An excerpt for the garbled.pdf test fixture (a PDF whose font encoding turns Czech diacritics
into mojibake), with no stage configured:
{
"pages": [{
"page": "1", "page_index": 1,
"text_layer": "garbled",
"needs_ocr": true,
"needs_ocr_reason": "embedded text layer does not decode: 3 of 3 lines carry CP1250 bytes read as CP1252 (mojibake diacritics), β¦",
"category": null,
"quality": {"source": "digital-convert", "score": 0.9356, "band": "Clear",
"lines_by_category": {"Garbage": 3, "decoded": 0}, "lang": null, "lines_scored": 3},
"route": "ocr",
"route_reason": "the text layer does not decode; page type unknown (page-classification not run): default to ATR",
"layout": {"canvas": {"width": 612.0, "height": 792.0, "unit": "pt"}, "lines": 3, "blocks": 1,
"tables": 0, "headings": 0, "columns": 1, "images": 0, "vector_paths": 0,
"regions": {"page_header": 0, "page_footer": 0, "footnote": 0}},
"text": "ZprΓ‘va o sondΓ¬ Γ¨Γslo 3.\nβ¦"
}],
"summary": {"pages": 1, "routes": {"nlp": 0, "ocr": 1, "htr": 0, "none": 0},
"needs_ocr_pages": [1], "reacquire_pages": [1], "document_route": "ocr"},
"stages": [
{"stage": "page-classification", "status": "not_configured", "detail": "PAGE_CLASSIFICATION_URL is not set", "record_adopted": false},
{"stage": "ocr-postprocess", "status": "not_configured", "detail": "OCR_POSTPROCESS_URL is not set", "record_adopted": false}
]
}-
Text layer (
text_layer), from the converter; the record carries the same value inpages[].text_layer:digital: the layer decodes;garbled: a layer that does not decode;ocr: a prior OCR run;none: no text layer;blank: an empty page of a format without page images.
-
Page type (
category): page-classification's label, confidence and top-N. It is asked about theneeds_ocrpages by default, or about every page withclassify_pages=all. PDF only. -
Readability (
quality):- with ocr-postprocess configured: its line-quality model (Clear / Noisy / Trash per line, the page band, the language);
- without it: the converter's decode check only. Its
score/bandmeasure how the characters decode, not how the text reads.
-
Route (
route), deterministic, withroute_reason:The page Route its text layer decodes nlp, unless ocr-postprocess calls more thanROUTE_TRASH_SHARE(0.5) of its linesTrashβocrno usable text, handwritten ( TEXT_HW,LINE_HW)htrno usable text, printed, typed or tabular ocrno usable text, graphical only ( DRAW,PHOTO)noneno usable text, page type unknown ocrblank noneThe categories are read through the page-category facets of
atrium_vocab.COLLECTIONS, so a category added to the vocabulary registry is routed without a code change. -
Layout and text (
layout,text): the converter's counts per page, and the page's text in reading order (include_text=falseleaves the text out). -
Stages (
stages[]), one entry per stage:- the
status:ok,partial,not_configured,not_requested,skipped,unavailable,errororrejected; - how long the call took, and the stage's version;
- whether its record was adopted.
A stage that is down or slow is reported, never a failed request.
- the
-
The stages' record writes. Each stage may write only its own fields:
- page-classification: the page categories;
- ocr-postprocess: the line and page quality.
digital-convert adopts the returned record only when the converter's own part is intact (a guard in
service/stages.py); otherwise the stage isrejectedand its answer goes intopages[]only. -
Trying it before the stages are deployed.
tools/stage_stub.pyis a stand-in for both stages:docker compose --profile stub up, orpython tools/stage_stub.py --port 8090withPAGE_CLASSIFICATION_URL=OCR_POSTPROCESS_URL=http://127.0.0.1:8090. It is never part of an image.
/reformat with markdown=true returns the record rendered as annotated Markdown. The same
renderer is available on the command line, and the TEITOK/ALTO renderer sits beside it:
python3 api_util/json_to_md.py CTX000000001.document.json --detail standard
python3 api_util/xml_to_md.py CTX000000001.teitok.xml --format layout --detail minimal
python3 api_util/doc_to_visual_md.py report.pdf --output report.md # convert + render in one step- The format. The Markdown is page-sectioned (
## Page N). The visual-layout cues ride in HTML comments (DOC_META,BBOX,PAGE_BREAK,NEEDS_OCR,HEADER_*/FOOTER_*, headings, GFM tables, footnotes). The full taxonomy isCUE_SCHEMAinapi_util/layout_md.pyπ. - The one model-facing representation. Annotated Markdown is the single general input a language
model is given (digital-convert#3). PDF, DOCX, TEITOK, PAGE XML and ALTO stay upstream, as sources;
nothing downstream takes raw HTML or XML as a prompt. The record route and the TEITOK/ALTO route emit
the same cue vocabulary and the same page labels. What the model said (
enrichment) is never read back into the Markdown, and entities, keywords and enrichment are checked on the record, not in the Markdown.tests/test_md_stress.pyholds both routes to this at every detail profile: each page and line once and in order, the cues each profile promises, the grouped structure, noenrichment, and the same output on every run and hash seed. - Who uses it. The whole-document runs of the keyword clients read
.mdfiles (atrium-keyword-extract'sopenrouter_client.pyandollama_client.py); its service works line by line on the record and does not need it. How the renderer is shared with that repository (vendored with a pin) is still to be settled (atrium-project#72 B). - One source. The rendering is a pure function of the record, so the Markdown cannot differ from the JSON.
Three cue profiles, the values of the record's regenerable.markdown.detail
(hub #70, item 1):
- the text lines are the same in all three; a lighter profile only drops cues;
- the cue sets nest (minimal β standard β full);
fullis the default.
| Cue | full |
standard |
minimal |
|---|---|---|---|
# doc, ## Page, PAGE_BREAK, NEEDS_OCR, headings, footnotes, GFM tables, HEADER_*/FOOTER_*, figure placeholders |
β | β | β |
OCR provenance, DOC_META, whole-line **bold**/*italic* |
β | β | β |
BBOX |
per line | one per block (a group_id run; a table keeps its own; ungrouped lines none) |
β |
LAYOUT_MARGIN (canvas minus the body-line union; record route, pages with a canvas and boxes) |
β | β | β |
What each profile costs, from python3 scripts/detail_budget.py (characters; βtokens = chars / 4):
| Input | full | standard | minimal |
|---|---|---|---|
hub E2E scan CTX192100040.alto.xml (xml_to_md layout) |
32 234 (β8 058) | 21 102 (β35 %) | 15 535 (β52 %) |
hub E2E scan CTX192601143.alto.xml (xml_to_md layout) |
26 614 (β6 653) | 12 555 (β53 %) | 9 536 (β64 %) |
CTX000000002.teitok.xml writer sample (xml_to_md layout) |
785 | 529 (β33 %) | 385 (β51 %) |
two_column.pdf fixture (JSON route) |
1 293 | 907 (β30 %) | 625 (β52 %) |
enrichable.pdf fixture (JSON route) |
596 | 354 (β41 %) | 282 (β53 %) |
rich.docx fixture (JSON route; DOCX has no boxes) |
632 | 632 (0 %) | 628 (β1 %) |
- The recipe. The record's
regenerable.markdownrecipe (json_to_md@1.1) is written only when the record can actually be rendered. - Unemitted cues. Cues in the catalogue that no profile emits (
INDENT,LAYOUT_COLUMN,FONT,STYLE,WATERMARK, strike, underline, alignment) are listed, each with its reason, inlayout_md.RESERVEDπ. - The deprecated direct converters.
api_util/pdf_to_md.pyanddocx_to_md.pyare reached only throughdoc_to_visual_md.py --legacy/--ocr, for A/B checks, and renderfullonly.
TEITOK input. xml_to_md.py reads a TEITOK (*.teitok.xml) or raw ALTO document.
- The TEITOK files come from atrium-nlp-enrich (format 2,
teitok-2) or from flexiconv. - The line-level reader
api_util/teitok_read.pyπ is vendored byte-identical from atrium-nlp-enrich, together withbbox_scale.pyandflexiconv_convert.py, and pinned bytests/test_vendored_teitok_parity.py. - The format itself is described in nlp-enrich's README, section TEITOK XML β Unified Output Format.
Every setting is an environment variable. .env.example π is the complete ledger
(tests/test_env_contract.py keeps it complete).
| Variable | Default | Meaning |
|---|---|---|
MAX_UPLOAD_MB, MAX_PAGES, MAX_CONCURRENT_JOBS |
50, 2000, 2 | service limits |
OCR_LAYER_DOCUMENT_SHARE |
0.5 | at or above this share of prior-OCR pages a PDF is refused (ocr_text_layer) |
ROUTE_TRASH_SHARE |
0.5 | /describe: a decoding page with more Trash lines than this goes to ocr |
LIBREOFFICE_BIN, LIBREOFFICE_TIMEOUT_S |
soffice, 120 s |
DOC/XLS conversion |
PAGE_CLASSIFICATION_URL, OCR_POSTPROCESS_URL |
unset | /describe's stages; unset β not_configured |
STAGE_TIMEOUT_S |
120 s | per stage call; over it the stage is unavailable |
PAGE_CLASSIFICATION_VERSION, PAGE_CLASSIFICATION_TOPN |
all (the ensemble), 3 |
the model version and top-N /describe asks page-classification for |
DOCLING_ARTIFACTS_PATH |
unset | Docling's model directory (--engine docling) |
The full table of limits, with what happens over each one, is in
service/README.md Β§ Limits. GET /info reports the value in force.
| Target | Image | What it is |
|---|---|---|
api |
ghcr.io/ufal/atrium-digital-convert-api |
the production image: /reformat, /describe (light engine + LibreOffice + FastAPI) |
digital |
ghcr.io/ufal/atrium-digital-convert-digital |
the converter's command line (published, never pinned) |
digital-docling |
built locally | the heavy PDF engine, with Docling's models downloaded at build time |
The tag is the release without its leading v (1.1.0-beta for v1.1.0-beta); the target is
part of the image name.
docker compose --profile api up # the service on :8000
docker compose --profile stub up # the service + the stage stand-ins
docker compose --profile digital run --rm digital-convert-digital \
/data/report.pdf --document-json-out /data/report.document.json
docker build --target digital-docling -t atrium-digital-convert-digital-docling .- Image contents.
.github/production-image.jsonπ declares the first-party files of theapiimage. The hub'stools/ci/image_closure.pychecks it on every push; research tools, tests and the stage stub stay out (.dockerignore). - Image size. LibreOffice (writer and calc,
-nogui) adds roughly 300 MB todigitalandapi.
Note
Docker on Linux: run as yourself. ./data belongs to you, while the images run as uid 10001
by default.
- With compose:
docker-compose.yamlruns every service asuser: "${ATRIUM_UID:-10001}:0", so put your uid in.envonce (echo "ATRIUM_UID=$(id -u)" >> .env). - With
docker run -v "$PWD:/data": pass--user "$(id -u):0". - Docker Desktop (macOS, Windows) needs neither. (atrium-project#69)
- The service. Every
/reformatand/describeresponse carriesparadata: the run's RO-CrateCreateAction(atrium_rocrate), programdigital-convert. It names the run id, what the call read (the file, with its digest, and the seed record when one was sent) and the record blocks it wrote. - Stage paradata.
/describealso returns each stage's own paradata instages[]. - The command line.
--paradata-dir DIRwritesYYMMDD-HHmmss_digital-convert.json(atrium_paradata.pyπ).
Licence of a record. provenance.license is the most restrictive licence among the components
the run actually used, as declared in para_config.txt π:
- The light engine and the shared reader are MIT, BSD-3-Clause or Apache-2.0, so a record says MIT.
- The Docling engine adds TableFormer's CDLA-Permissive-2.0, which the shared licence table ranks with MIT, so a Docling record also says MIT.
- LibreOffice (MPL-2.0) is declared but never logged. It converts the container format of
a text it does not author, and AMΔR accepted it on exactly that condition (#4). The conversion is
reported beside the record instead (
reader.conversionin the response). - PyMuPDF (AGPL-3.0) is deliberately not used; see
requirements_digital.txt.
The record follows atrium_document.schema.json π, schema version
1.0, frozen as the hub tag
doc-schema-v1
(ufal/atrium-project@544298b).
- tests/test_schema_freeze.py π checks this repository's schema against the frozen copy beside it, atrium_document.schema.doc-schema-v1.json π.
- The schema files,
atrium_document.pyand the other shared modules (atrium_service.py,atrium_openapi.py,atrium_limits.py,atrium_vocab.py,atrium_rocrate.py, β¦) are vendored from the hub. Thepara-driftworkflow keeps them byte-identical, so they are never edited here; the hub'sscripts/revendor_shared.shrefreshes them.
What may change after the freeze is in the hub's Freeze & conformance.
The evaluation harness of hub issue #22 (see #3) compares out-of-the-box VLM/OCR models with the legacy ABBYY/ALTO pipeline, scored per quality tier on an in-domain gold set.
- Scripts.
sample_stratify.py,bench_compare.pyandeval_metrics.pyare torch-free. - Where they run. They are research tools, run from a checkout and left out of every image
(
.dockerignore, atrium-project#72).
python sample_stratify.py --page-stats samples_page_stats.csv --n 200 --output docu_sample_manifest.csv
python bench_compare.py --manifest docu_sample_manifest.csv --gold data/gold \
--pred alto=../atrium-ocr-postprocess/data_samples/PAGE_TXT \
layoutreader=../atrium-ocr-postprocess/data_samples/PAGE_TXT_LR \
--split test --output-dir bench_results- Sampling.
sample_stratify.pybuckets pages by OCR quality (clean/degraded/hard/text_poor) and writes an annotation manifest with a deterministic 80/10/10 split. - Gold data. One UTF-8 transcription per manifest page
(
gold/<doc>/<doc>-<page>.txt, optionally.entities.tsv). - Comparison.
bench_compare.pywritespage_scores.csv,aggregate_scores.csvandreport.md(CER, WER, NED, optional entity P/R/F1). The output is byte-identical across reruns. - Config. Both scripts also read an INI config (
--config config_docu.txt, sections[STRATIFY]and[BENCHMARK]).
pip install -r requirements.txt -r requirements_digital.txt -r requirements-test.txt
pytest -q # the whole suite (no LibreOffice needed: DOC/XLS use a fake soffice)
ruff check . && ruff format --check .
python tests/fixtures/digital/make_fixtures.py --verify # the generated PDF/DOCX fixtures are current
python atrium_openapi.py export --app service.api:app --out service/openapi.json # after an API change
python atrium_openapi.py check --app service.api:app --spec service/openapi.json
python ../atrium-project/tools/ci/image_closure.py --repo-root . --worktree # the image's file listHow to contribute, the release procedure and the history are in
CONTRIBUTING.md π. The design record of every issue is in
agent_dev_logs/ (digests and plans).
For support write to: lutsai.k@gmail.com responsible for this GitHub repository 1 π
- Developed by UFAL 2 π₯
- Funded by ATRIUM 3 π°
- Shared by ATRIUM 3 & UFAL 2 π
- Pipeline partners: AMΔR (the Archaeological Map of the Czech Republic) and the ATRIUM hub 4
- Frameworks used:
- pdfplumber / pdfminer.six and pypdfium2 (PDF text, geometry, page labels)
- python-docx and lxml (DOCX, and the ODF/OOXML parts in the shared reader)
- Docling 5 (the opt-in layout and table engine)
- LibreOffice 6 (legacy DOC/XLS conversion, headless)
- FastAPI + uvicorn (the service)
- UFAL flexiconv 7 (optional TEITOK conversion adapter)
Β©οΈ 2026 UFAL & ATRIUM