Poster at NeurIPS 2026 TAE Workshop

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

An owner-disjoint audit of near-tied leaderboard orderings
with residual family DIF and matched-random controls

Qiaoyuan Zheng·Yiqu Yang

ETH Zurich

TAE (Trust-AI-Eval): Can We Trust AI Evaluation?

The idea

A sub-one-point gap
is not self-interpreting.

Three panels. A: residual family DIF is estimated from the response matrix with a family-label-free spectral MIRT term held fixed; low-DIF items behave similarly across families, high-DIF items do not. B: the full benchmark is recomposed into a low-DIF subtest and a matched-random subtest of the same size, balanced by item group and easiness. C: under low-DIF scoring a near-tied pair of models from different families reverses order while another near tie is preserved.
Overview · The audit in three steps Residual DIF estimation, blueprint-preserving recomposition, and ranking comparison against matched-random subtests. Model letters and scores in panel C are illustrative. Original vector PDF ↗
5benchmarks
43,144items with item-level responses
5model families
1,000matched-random subtests per benchmark

Leaderboards often separate language models by less than one percentage point. Such a gap only supports “model A is better than model B” if the ordering does not hinge on which items the benchmark happens to include. We test whether near-tied cross-family orderings survive when a benchmark is recomposed toward items whose difficulty does not depend on model family—beyond what an equally short random subtest would do.

  1. Estimate residual family DIF

    A family-label-free spectral approximation to multidimensional IRT absorbs the ability structure. The remaining item-by-family effects measure residual differential item functioning across Qwen, Llama, Gemma, Mistral/Mixtral, and Phi.

  2. Recompose without leakage

    In owner-disjoint folds, one owner half selects low-DIF anchor items. Frozen weights score models in the other half and restore every item-group × easiness cell to its original mass.

  3. Compare against matched random

    For cross-family pairs within one point, we count strict rank reversals and compare them with 1,000 subtests of the same length and composition that ignore DIF. The excess isolates targeted composition sensitivity.

The evidence

Global rankings hold.
Near ties do not.

Full-benchmark and low-DIF rankings remain strongly correlated, and only 2.5–4.2% of all cross-family comparisons reverse. Among pairs initially within one percentage point, however, four of five benchmarks reverse far more often than matched-random subtests.

All numbers on this page come from arXiv:2609.00482 and are reproducible from the frozen tables in the repository. This is not a live leaderboard.

Near-tie reversals

30.9–47.1%

of cross-family pairs within one point reverse order under low-DIF scoring in MMLU-Pro, BBH, MMLU, and HellaSwag.

Beyond matched random

+16.9–28.6 pp

excess over the median of 1,000 matched-random subtests (all one-sided p = .001). WinoGrande shows no reliable excess (−0.9 points, p = .689).

Primary near-tie results with 50% residual low-DIF anchors. K: selected spectral-MIRT dimensions in the two folds. τb: mean fold-specific Kendall correlation between full-benchmark and low-DIF rankings. Close pairs: different-owner, different-family pairs within one percentage point.
BenchmarkKKendall τbClose pairsLow-DIF reversalsMatched randomExcess (pp)p
MMLU-Pro64/32.92430847.1%18.5%+28.6.001
BBH128/128.9152,22642.1%25.2%+16.9.001
MMLU128/128.9303,23540.7%16.3%+24.4.001
HellaSwag16/64.9484,53330.9%11.4%+19.5.001
WinoGrande8/8.9005,57433.8%34.7%−0.9.689
Excess near-tie reversals across maximum full-score gaps from 0.25 to 5 percentage points. MMLU-Pro, BBH, MMLU, and HellaSwag stay well above zero; WinoGrande stays near zero.
Figure 1 · Excess persists across wider score bands At a five-point gap, the four primary-positive benchmarks still exceed matched-random controls by 5.2–19.1 points; WinoGrande shows no significant excess at any threshold. Thresholds other than the pre-specified one point are descriptive. Original vector PDF ↗

Robustness

It is not about who is in the pool.
No family always wins.

We re-ran the full audit under nine population perturbations: owner caps of 1 and 3, score-blind checkpoint selection, a conservative lineage filter, and leaving out each family in turn. Each variant refits the response directions, residual DIF, anchors, and matched controls.

Excess reversal for each benchmark under the baseline and nine population specifications. MMLU-Pro, BBH, MMLU, and HellaSwag remain positive and significant in every specification; WinoGrande is significant only when Mistral/Mixtral is omitted.
Figure 2 · Population robustness Diamonds mark the primary specification; filled markers denote one-sided matched-random p ≤ .05. Original vector PDF ↗

The effect is consistent; the winner is not. MMLU-Pro, BBH, MMLU, and HellaSwag show positive excess with p ≤ .05 in every specification. Yet every family changes the direction of its mean score shift at least once across benchmarks, so the shifts should not be read as intrinsic family rankings.

Interpretation

The signatures replicate.
A simple story does not.

Owner-disjoint replication

31.5–38.5%

top-20% high-DIF overlap across disjoint owner halves, against a 20% independence baseline; median family-wise ρ = .308–.589 (all permutation p = .002).

Blinded content audit

0 of 10 axes

meet the pre-specified dual-annotator confirmatory criteria in 250 matched high-DIF/control pairs, despite good agreement (κ = .715).

Paired prevalence differences between stable high-DIF items and matched controls on ten binary content axes for two blinded annotators. No axis exceeds the eight-point gate for both annotators.
Figure 3 · Blinded content audit (appendix) The largest effect, annotator A’s −9.6-point spatial/temporal difference (q = .106), is neither significant after correction nor replicated by annotator B. Original vector PDF ↗

Residual DIF is a diagnostic conditional on a linear, family-blind response model. It does not establish unfairness, contamination, or a causal family mechanism, and the low-DIF score is not proposed as a more correct replacement score.

Use it

Verify the paper.
Audit your own benchmark.

The repository contains the implementation, frozen protocols, aggregate results, and tests. Raw benchmark text and response matrices are not redistributed; version-pinned preparation tools rebuild them from the public upstream releases.

  1. Verify offline

    Run the tests and release audit against the frozen results. No data download or model calls are needed.

  2. Audit your benchmark

    Provide a binary item-by-model response matrix with model families, owners, and item groups; the CLI reports near-tie reversals and their matched-random excess.

  3. Recompute the paper

    Download and checksum the RouterEval and MMLU-Pro inputs, then rerun every stage with the staged reproduction command.

Install and verify (Python 3.12)
git clone https://github.com/qiaoyuan667/near-tie-robustness.git
cd near-tie-robustness
python3.12 -m venv .venv && source .venv/bin/activate
python -m pip install -r requirements-pinned.txt
python -m pip install -e . --no-deps
make verify
Audit your own benchmark
family-dif-audit your-benchmark.npz --output your-benchmark-audit

A nonsignificant result does not certify that a benchmark’s rankings are robust in general. Keep the pinned numerical environment: exact-tie counts are sensitive to floating-point library changes.

Build on this work

Cite the paper.

If this audit or its code supports your research, please cite:

BibTeX
Download .bib
@article{zheng2026near,
  title={Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?},
  author={Zheng, Qiaoyuan and Yang, Yiqu},
  journal={arXiv preprint arXiv:2609.00482},
  year={2026}
}

Scope and reuse

Family labels are observational and inferred from model identifiers. The five benchmarks are static, mostly closed-form, and drawn from one response collection; MMLU and MMLU-Pro share lineage. The one-point band is operational, and DIF does not imply unfairness.

Code, documentation, figures, and released result files: MIT, to the extent the authors hold rights in them. Third-party benchmark data keep their upstream licenses. All plots on this page come from the paper’s released figures.