Lemmatised BM25 adds ≥ 10 points of Recall@10 on Czech questions.
D1 → H1 delta on the Czech and cross-lingual slices.
dev split. Real numbers from the tuning set, not the held-out test split; retrieval only, no answer layer has run yet.
Each step adds one change to the previous config. Deltas are paired over the same questions with a 1,000× bootstrap; a step whose 95% CI includes zero is labelled n.s. With 100 dev questions many steps will be — that is a finding, not a failure.
| Step | Change | Recall@10 | Δ vs previous [95% CI] | Significance |
|---|---|---|---|---|
| D1 | Naive dense (bge-m3, fixed-512) | 52.4 [43.7–61.4] | baseline | |
| H1L | + lemmatised BM25 via RRF, structure chunks | 55.9 [47.1–64.6] | +3.5 pp [−5.9, +12.2] | n.s. |
| T1L | + table-row and XLSX row chunks | 50.6 [42.0–58.8] | −5.4 pp [−10.6, −0.6] | CI excludes 0 |
Bootstrap CIs here are per comparison; the harness's paired permutation test with Holm correction is the significance of record.
BM25 over raw tokens (B1) against the same BM25 over MorphoDiTa lemmas (B2). The pair is outside the waterfalls, which change one thing per step along another path.
| Questions | n | Δ B1 → B2 [95% CI] | Significance |
|---|---|---|---|
| All answerable | 90 | −2.4 pp [−7.4, +2.2] | n.s. |
| Czech only | 49 | −4.4 pp [−12.9, +4.8] | n.s. |
Lemmatised BM25 adds ≥ 10 points of Recall@10 on Czech questions.
D1 → H1 delta on the Czech and cross-lingual slices.
Rerankers hurt table-row chunks.
R1 vs H1 on numeric questions; rerank-stage misses in /failures.
Most scan failures come from PARSE, not retrieval.
Stage split of misses on questions tagged scan.
Long context matches RAG on accuracy but costs ≥ 10× more per query and cites less precisely.
LC vs T1 accuracy, $/1k queries and citation precision.
Contextual retrieval helps version and multi-hop questions the most.
R1 → C1 delta by category; needs the full test split.
Failure classes, the fix for each, and the re-measured delta. The rows below are candidates that have not been re-measured yet; each delta is filled in only once its fix has run.
| Failure class | Fix | Before → after |
|---|---|---|
| Split-table header: row 14 chunked away from its column headers | Table-row chunks repeat the header path and unit | pending |
| mil. vs tis. Kč normalisation in numeric answers | Unit-aware Czech number normaliser in the numbers-verified check | pending |