
Biased LLM judges disable skill retirement in self-improving agents
A new study reveals a critical failure mode in self-improving AI agents that use LLM judges to evaluate their own skills. When reward signals are biased, the mechanism that normally prevents skill degradation silently breaks down, allowing poor capabilities to persist in the agent's library. The researchers isolate this causal failure through controlled experiments on code generation and report writing, showing that asymmetric bias (false passes) is far more damaging than symmetric noise. This finding matters for anyone building reference-free evaluation systems or autonomous agents that learn from their own feedback loops, as it exposes a hidden vulnerability in self-curation architectures that current safety assumptions don't account for.62























