Poster at NeurIPS 2026 TAE Workshop
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
An owner-disjoint audit of near-tied leaderboard orderings
with residual family DIF and matched-random controls
ETH Zurich
TAE (Trust-AI-Eval): Can We Trust AI Evaluation?
The idea
A sub-one-point gap
is not self-interpreting.
Leaderboards often separate language models by less than one percentage point. Such a gap only supports “model A is better than model B” if the ordering does not hinge on which items the benchmark happens to include. We test whether near-tied cross-family orderings survive when a benchmark is recomposed toward items whose difficulty does not depend on model family—beyond what an equally short random subtest would do.
Estimate residual family DIF
A family-label-free spectral approximation to multidimensional IRT absorbs the ability structure. The remaining item-by-family effects measure residual differential item functioning across Qwen, Llama, Gemma, Mistral/Mixtral, and Phi.
Recompose without leakage
In owner-disjoint folds, one owner half selects low-DIF anchor items. Frozen weights score models in the other half and restore every item-group × easiness cell to its original mass.
Compare against matched random
For cross-family pairs within one point, we count strict rank reversals and compare them with 1,000 subtests of the same length and composition that ignore DIF. The excess isolates targeted composition sensitivity.
The evidence
Global rankings hold.
Near ties do not.
Full-benchmark and low-DIF rankings remain strongly correlated, and only 2.5–4.2% of all cross-family comparisons reverse. Among pairs initially within one percentage point, however, four of five benchmarks reverse far more often than matched-random subtests.
All numbers on this page come from arXiv:2609.00482 and are reproducible from the frozen tables in the repository. This is not a live leaderboard.
Near-tie reversals
30.9–47.1%
of cross-family pairs within one point reverse order under low-DIF scoring in MMLU-Pro, BBH, MMLU, and HellaSwag.
Beyond matched random
+16.9–28.6 pp
excess over the median of 1,000 matched-random subtests (all one-sided p = .001). WinoGrande shows no reliable excess (−0.9 points, p = .689).
| Benchmark | K | Kendall τb | Close pairs | Low-DIF reversals | Matched random | Excess (pp) | p |
|---|---|---|---|---|---|---|---|
| MMLU-Pro | 64/32 | .924 | 308 | 47.1% | 18.5% | +28.6 | .001 |
| BBH | 128/128 | .915 | 2,226 | 42.1% | 25.2% | +16.9 | .001 |
| MMLU | 128/128 | .930 | 3,235 | 40.7% | 16.3% | +24.4 | .001 |
| HellaSwag | 16/64 | .948 | 4,533 | 30.9% | 11.4% | +19.5 | .001 |
| WinoGrande | 8/8 | .900 | 5,574 | 33.8% | 34.7% | −0.9 | .689 |
Robustness
It is not about who is in the pool.
No family always wins.
We re-ran the full audit under nine population perturbations: owner caps of 1 and 3, score-blind checkpoint selection, a conservative lineage filter, and leaving out each family in turn. Each variant refits the response directions, residual DIF, anchors, and matched controls.
The effect is consistent; the winner is not. MMLU-Pro, BBH, MMLU, and HellaSwag show positive excess with p ≤ .05 in every specification. Yet every family changes the direction of its mean score shift at least once across benchmarks, so the shifts should not be read as intrinsic family rankings.
Interpretation
The signatures replicate.
A simple story does not.
Owner-disjoint replication
31.5–38.5%
top-20% high-DIF overlap across disjoint owner halves, against a 20% independence baseline; median family-wise ρ = .308–.589 (all permutation p = .002).
Blinded content audit
0 of 10 axes
meet the pre-specified dual-annotator confirmatory criteria in 250 matched high-DIF/control pairs, despite good agreement (κ = .715).
Residual DIF is a diagnostic conditional on a linear, family-blind response model. It does not establish unfairness, contamination, or a causal family mechanism, and the low-DIF score is not proposed as a more correct replacement score.
Use it
Verify the paper.
Audit your own benchmark.
The repository contains the implementation, frozen protocols, aggregate results, and tests. Raw benchmark text and response matrices are not redistributed; version-pinned preparation tools rebuild them from the public upstream releases.
Verify offline
Run the tests and release audit against the frozen results. No data download or model calls are needed.
Audit your benchmark
Provide a binary item-by-model response matrix with model families, owners, and item groups; the CLI reports near-tie reversals and their matched-random excess.
Recompute the paper
Download and checksum the RouterEval and MMLU-Pro inputs, then rerun every stage with the staged reproduction command.
git clone https://github.com/qiaoyuan667/near-tie-robustness.git
cd near-tie-robustness
python3.12 -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements-pinned.txt
python -m pip install -e . --no-deps
make verifyfamily-dif-audit your-benchmark.npz --output your-benchmark-auditA nonsignificant result does not certify that a benchmark’s rankings are robust in general. Keep the pinned numerical environment: exact-tie counts are sensitive to floating-point library changes.
Build on this work
Cite the paper.
If this audit or its code supports your research, please cite:
@article{zheng2026near,
title={Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?},
author={Zheng, Qiaoyuan and Yang, Yiqu},
journal={arXiv preprint arXiv:2609.00482},
year={2026}
}Scope and reuse
Family labels are observational and inferred from model identifiers. The five benchmarks are static, mostly closed-form, and drawn from one response collection; MMLU and MMLU-Pro share lineage. The one-point band is operational, and DIF does not imply unfairness.
Code, documentation, figures, and released result files: MIT, to the extent the authors hold rights in them. Third-party benchmark data keep their upstream licenses. All plots on this page come from the paper’s released figures.