首页 > AI前沿 > JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

arXiv自然语言 2026-10-05 23:41 4 阅读 查看原文

When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression.

We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score.

A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding.

Performance Improvements

On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by +0.079.

Applying the same configuration to six additional cells yields a +0.0393 mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank p=0.0091.

Validation-Based Analyses

Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.