Mechanistic analysis reveals how LLM judges evaluate text quality
Researchers have opened the black box of LLM-based evaluators by mechanistically analyzing how Themis and Prometheus assign quality ratings to generated text. Using controlled perturbations across readability and adequacy dimensions, combined with causal tracing and attention analysis, the work reveals that both models execute a coherent two-stage evaluation pipeline. This matters because LLM judges now drive both automated scoring and training signals across the industry, yet their internal decision-making remained opaque. Understanding these mechanisms is critical for practitioners deploying evaluators in production and for researchers building more reliable NLG assessment systems.62