What actually works for retrieval over a messy Czech/English sell-side data room.
A deliberately messy synthetic data room — scanned Czech statutory accounts, merged-cell management accounts, draft/final/signed versions, planted conflicts and decoys — whose gold evidence is known down to the page. 5 pipelines are scored on whether they retrieve the evidence each question requires, with latency per query, and every miss is traced to the stage that lost it.
This run is retrieval only. Answer accuracy, conflict handling, abstention and cost need the answer layer, which has not run yet, so the site does not report them.
Evidence recall@10
- D1 · naive
- 52% 95% CI 43–61
- H1L · highest of 5
- 56% 95% CI 46–65
H1L has the highest evidence recall@10 of 5 configs, but the 95% intervals overlap: not a demonstrated gain. The paired test is on the ablation page.
See ablation →Where misses are lost
- H1L · at recall
- 42 of 52 81% of misses · 10 at parse
First stage that lost the gold, dev split. Retrieval only: unanswerable questions count as no miss.
See failures →Required evidence parsed
- answerable questions
- 80 of 90 89%
Questions whose every required evidence item survives parsing (OCR for scans) in H1L's corpus; the rest are lost at parse.
See failures →Query latency, p50
- H1L · best
- 17 ms p95 19 ms
- B1 · fastest
- 0.1 ms p95 0.1 ms
Per query over all retrieval stages, indexing excluded; calls served from the cache are not timed.
See results →Findings
Computed from the dev-split run; the held-out test split may disagree.B1 and B2 retrieve none of the cross-lingual evidence
On the 12 cross-lingual dev questions (asked in English about Czech documents), D1 reaches recall@10 of 58%, while B1 and B2 score 0%. Descriptive: a slice of 12 questions, no interval.
Config × category heatmap →2Most misses are lost at retrieval, not parsing
Across 5 configs and 90 answerable dev questions, 42–55 are lost at recall (a required evidence group not retrieved) and 10 at parse (required evidence missing from the parser output).
Failure funnel →3Czech lemmatisation shows no reliable gain
BM25 over MorphoDiTa lemmas (B2) vs raw tokens (B1): recall@10 −2.4 pp [−7.4, +2.2], paired over 90 dev questions; the interval includes zero. On the 49 Czech questions alone: −4.4 pp [−12.9, +4.8].
Pairwise comparison, B1 → B2 →Start here
- Ask — one question, two configs side by side, their top chunks against the gold evidence.
- Failures — click a miss, see the gold region and where it ranked.
- Ablation — each step's delta with its CI, including what did not help.
- Czech scan — the case study on a public Czech filing, and its status.
Reproducibility
- Gold evidence is known by construction, down to page and bounding box.
- Every model call is cached by request hash, so a rerun reproduces every retrieval metric exactly; only latency is measured afresh.
- These are dev-split numbers, from the tuning set; the test split stays held out until the configs are frozen.
- Results generated 2026-10-01T05:20:25+00:00.