Amazon research uses Ising models to assess LLM judge bias

A seemingly strong consensus among large language model (LLM) judges may be misleading, according to new research from Krishna Balasubramanian and Sasha Podkopaev. The team discovered that correlated errors can inflate apparent agreement; for example, strong validation is not necessarily indicated, International Conference on Machine Learning says.

Their method, using Ising models to assess correlations between judge outputs and adjust scores, outperformed the best-performing baseline, a panel of judges weighted according to historical accuracy, by 9% to 14% on standard metrics. This dependence-aware aggregation treats a judge panel as a network, learning both individual skill and relationships between judges.

Ising Models Assess LLM Judge Correlation for Reliable Evaluation

The team reports improvements of between 9% to 14% on standard metrics across three tasks by accounting for these correlations. This refinement surpasses the accuracy of historically accurate weighted panels, which treat judge errors as independent events. The work, detailed in “Dependence-aware label aggregation for LLM-as-a-judge via Ising models,” presented at this year’s International Conference on Machine Learning, introduces a method to assess these relationships and adjust aggregate scores accordingly.

The researchers evaluated their approach on relevance classification, toxicity classification, and summarization assessment, utilizing panels of 10 judge models operating without randomness, ensuring consistent outputs for identical inputs. Weighted and uniform majority vote systems serve as useful baselines, but they operate under the assumption that incorrect answers stem from independent errors, a simplification often inaccurate for LLM judges.

Judges may share similar interpretations of evaluation rubrics or inherit biases from identical prompting examples, leading to correlated failures. This allows for identification of redundant judges and task-specific blind spots during audits, answering questions about whether added judges provide independent evidence or simply reinforce existing biases.

For example, on relevance classification, the dependence-aware model achieved 0.912 accuracy, compared with 0.820 for weighted majority vote and 0.804 for uniform majority vote. Similar gains were observed in toxicity classification, reaching 0.792 accuracy, compared with 0.694 and 0.695; and summarization assessment, attaining 0.806 accuracy, compared with 0.737 and 0.561, further demonstrating the benefits of modeling judge dependencies.

The learned network also provides insights into the nature of agreement; it can reveal if a task elicits broad consensus or if opinions cluster, informing the optimal composition of future judge panels and improving the reliability of LLM evaluation.

Unsupervised Learning Reveals Judge Relationships and Potential Biases

The detailed approach allows for identification of task-specific weaknesses in evaluation panels, going beyond simply measuring individual judge reliability. The team’s method uses the relationships between judges to pinpoint instances where shared blind spots might skew results, offering a more nuanced audit trail than traditional methods. This is achieved by modeling the panel as a network, where the agreement between any two judges isn’t solely determined by their individual skill, but also by their correlation, or lack thereof, with each other, according to International Conference on Machine Learning.

Ising models, a statistical tool for representing dependencies between binary variables, form the core of this new technique; they map the complex interactions within the judge panel. Unlike weighted majority voting, which assigns scores based on individual judge performance, this dependence-aware aggregation considers how judges influence one another.

The system operates without needing human-labeled data for training, instead inferring both judge reliability and the relationships between them from the judges’ outputs alone. “Each judge still has its own reliability profile, but pairs of judges can also have relationships,” the researchers explain, highlighting the method’s ability to capture subtle interactions. The researchers note that these improvements require sufficient evaluation data and a sizable judge panel to accurately estimate the relationships between judges.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of The Neuron

The Neuron

With a keen intuition for emerging technologies, The Neuron brings over 5 years of deep expertise to the AI conversation. Coming from roots in software engineering, they've witnessed firsthand the transformation from traditional computing paradigms to today's ML-powered landscape. Their hands-on experience implementing neural networks and deep learning systems for Fortune 500 companies has provided unique insights that few tech writers possess. From developing recommendation engines that drive billions in revenue to optimizing computer vision systems for manufacturing giants, The Neuron doesn't just write about machine learning—they've shaped its real-world applications across industries. Having built real systems that are used across the globe by millions of users, that deep technological bases helps me write about the technologies of the future and current. Whether that is AI or Quantum Computing.

Latest Posts by The Neuron: