dev split. Real numbers from the tuning set, not the held-out test split; retrieval only, no answer layer has run yet.

Ablation

Each step adds one change to the previous config. Deltas are paired over the same questions with a 1,000× bootstrap; a step whose 95% CI includes zero is labelled n.s. With 100 dev questions many steps will be — that is a finding, not a failure.

Questions

Waterfall: Recall@10

100 questions · dashed bar = whole-room long context
02040608010052.4D1n.s.H1L−5.4T1L
Waterfall steps with scores and paired deltas
StepChangeRecall@10Δ vs previous [95% CI]Significance
D1Naive dense (bge-m3, fixed-512)52.4 [43.7–61.4]baseline
H1L+ lemmatised BM25 via RRF, structure chunks55.9 [47.1–64.6]+3.5 pp [−5.9, +12.2]n.s.
T1L+ table-row and XLSX row chunks50.6 [42.0–58.8]−5.4 pp [−10.6, −0.6]CI excludes 0

Bootstrap CIs here are per comparison; the harness's paired permutation test with Holm correction is the significance of record.

Pairwise: Czech lemmatisation

recall@10 · answerable questions · not affected by the filters above

BM25 over raw tokens (B1) against the same BM25 over MorphoDiTa lemmas (B2). The pair is outside the waterfalls, which change one thing per step along another path.

Paired recall@10 delta from B1 to B2
QuestionsnΔ B1 → B2 [95% CI]Significance
All answerable90−2.4 pp [−7.4, +2.2]n.s.
Czech only49−4.4 pp [−12.9, +4.8]n.s.

Pre-registered hypotheses

committed before the first run · 5 of 5 verdicts pending the test-split run
H1 Pending

Lemmatised BM25 adds ≥ 10 points of Recall@10 on Czech questions.

D1 → H1 delta on the Czech and cross-lingual slices.

H2 Pending

Rerankers hurt table-row chunks.

R1 vs H1 on numeric questions; rerank-stage misses in /failures.

H3 Pending

Most scan failures come from PARSE, not retrieval.

Stage split of misses on questions tagged scan.

H4 Pending

Long context matches RAG on accuracy but costs ≥ 10× more per query and cites less precisely.

LC vs T1 accuracy, $/1k queries and citation precision.

H5 Pending

Contextual retrieval helps version and multi-hop questions the most.

R1 → C1 delta by category; needs the full test split.

Iteration log

Failure classes, the fix for each, and the re-measured delta. The rows below are candidates that have not been re-measured yet; each delta is filled in only once its fix has run.

Failure classFixBefore → after
Split-table header: row 14 chunked away from its column headersTable-row chunks repeat the header path and unitpending
mil. vs tis. Kč normalisation in numeric answersUnit-aware Czech number normaliser in the numbers-verified checkpending