dev split. Real numbers from the tuning set, not the held-out test split; retrieval only, no answer layer has run yet.

Method

How the numbers are made, how the project was built, and what to try next. Everything the site shows is precomputed from the harness's export and reproducible from its request cache.

The room

Project Beskyd is a fictional Czech manufacturer's sell-side data room, rendered from one seeded ledger of facts: 170 documents — PDFs (31 of them scans, 481 PDF pages in all), spreadsheets, Word files and e-mails — with 12 planted conflicts and 8 decoys. It carries 200 questions, split 100 dev / 100 test with 10 unanswerable questions on each side. This is room version 4; scores from earlier versions are not comparable.

Gold by construction

The generators record where each fact is drawn while they draw it: page and bounding box, or sheet and cell. When a page is rendered as a skewed scan, the stored affine transform moves its gold boxes too. Guard tests read the rendered room back as text and fail if a document contradicts a planted conflict or answers a question planted as unanswerable.

Audited gold

All 200 questions were audited against the rendered documents. The audit found 31 defects on 29 questions (ambiguous wording, missing or wrong evidence, wrong answers, mislabelled conflicts and decoys). Each was fixed at the source — ledger, renderers, templates, evidence selection — not in the generated files, and each defect class has a test that fails on the old generator.

What the metrics count

Each fact is required from its home document (the contract for contract terms, that year's statements for statutory figures); copies there share an evidence group, and finding any copy counts. Recall@10 is the share of a question's required groups in the top 10, MRR the mean of 1 / rank of the first chunk that satisfies a required group, nDCG is computed over groups. The same fact repeated elsewhere is a mention and earns no headline credit.

Alignment across parsers

Parsers cut pages differently, so block ids cannot be compared. A block holds gold when its box overlaps the gold box (IoU ≥ 0.3) and its text confirms the gold, or when at least 80% of the normalised gold tokens occur in it (case, diacritics and Czech number formats folded). A chunk holds gold only if the gold is intact in it.

Splits and statistics

Dev and test are split 50/50, stratified by category; questions that share a fact, a conflict or a computation land on the same side. Tuning happens on dev only. Every metric carries a 1,000-resample bootstrap 95% CI over questions. The harness tests each step with a paired permutation test and Holm correction; the ablation page shows paired bootstrap deltas. With 90 answerable questions per split, many steps are not significant, and the site says so.

Failure attribution

Each question is attributed to the first stage that lost a required evidence group: parse (not in the parser output), chunk (not intact in any chunk), recall (not retrieved; without a reranker, outside the final top k), rerank (dropped from the top k by the reranker). Generation and grounding stages apply once an answer layer runs.

Not measured yet

The published run is retrieval only. Answer accuracy, citation precision, conflict handling, abstention and cost need the answer layer. When it runs, numeric answers are to be checked deterministically after unit and Czech-format normalisation, and a numbers-verified check is to require every number in an answer to appear in a cited span. Until then the site shows none of these metrics.

How it was built

Typed contracts first: a frozen document model, a results bundle and a request cache, each with tests. The synthetic room, parsers, retrieval harness and this explorer were then built against those contracts, every change gated on lint, strict type checks and tests. The explorer was developed against a deterministic fixture before any real run existed, so swapping in real results is a file copy. The whole pipeline — room, OCR, corpora, retrieval runs and export — runs offline on a laptop with one make all, and every model call goes through the request cache, so a rerun reproduces every retrieval metric exactly.

What I'd try next

  1. Per-deal eval sets mined from reviewer accept/reject decisions, giving a regression eval on every deal.
  2. The same benchmark on a production stack: managed ranking and layout parsing vs open models, on Czech scans, sliced by messiness.
  3. Table-row and cell-level indexing for management-account spreadsheets, with cell citations.
  4. Conflict-aware answers: an authority tier per document type, explained-difference handling and the numbers-verified badge.
  5. Abstention that files an information request into the data-room coverage map.
  6. An eval regression gate in CI for changes to parsing, chunking or retrieval.