Publication
arXiv
Stage
Preprint
What we read
Summary of the abstract

What the paper reports

The authors tested whether repeated requests to model services produced consistent rankings. Their reported results fell below reliability thresholds they had set before collecting the data.

Why it matters

If an evaluator changes its judgments, a score can become difficult to interpret as a change in the system being tested.

The campaigns examined repeated requests and later replays. The authors report that several attempted adjustments did not resolve the inconsistency in the settings tested.

They propose testing the reliability of the judging system before relying on it as an evaluation instrument.

What this does not tell us

The abstract concerns external behaviour on shared infrastructure. It does not establish that every model or evaluation method is unreliable. This is a preprint.

Original sources · 1
  1. Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints ↗arXiv · 2026-09-03

Check the original paper for its authors, methods, version and access terms.