Modelwire
Subscribe

RefCon extracts agent memories without labels, scaling test-time compute

RefCon addresses a critical bottleneck in agent development: extracting usable knowledge from noisy real-world interactions without expensive retraining or human annotation. By combining iterative refinement with contrastive learning, the method enables agents to improve through test-time computation alone, delivering 16-35% gains across multiple benchmarks. This matters because it decouples agent capability scaling from labeled data scarcity, a persistent constraint in long-horizon reasoning tasks. The approach generalizes across frameworks, suggesting a reusable pattern for production systems where ground-truth feedback is unavailable or prohibitively costly.

Modelwire context

Explainer

RefCon's key contribution is decoupling the learning signal from labeled data entirely. Rather than requiring human annotation or RL reward functions, the method uses contrastive comparisons between agent trajectories to extract refinement signals at inference time, meaning capability gains happen without touching model weights or collecting new supervision.

This builds directly on the self-improvement pattern established by last week's Retrospection-Only Fine-Tuning work, which showed agents can improve through introspection alone. RefCon extends that insight by making the introspection mechanism explicit and learnable through contrastive memory extraction. It also complements the memory optimization work from late September (Learning What to Remember), which tackled credit assignment for what to store. Where that work solved the 'what' problem, RefCon solves the 'how to extract value from what you already have' problem, suggesting a two-layer approach to agent memory systems.

If RefCon's 16-35% gains replicate on BFCL-V3 when tested against held-out agent trajectories from production deployments (not just benchmark rollouts), that confirms the method generalizes beyond academic evaluation. If gains collapse when contrastive pairs come from different model families or task distributions, the approach is narrower than claimed.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsRefCon · AppWorld · BFCL-V3 · ACE · ReMe · ReasoningBank

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “RefCon: Iterative Refinement and Contrastive Memory Extraction for Context-Evolving Agent”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Related

ReCAP reduces memory overhead for long-running LLM agents

arXiv cs.CL·

Language agents improve through self-explanation without reinforcement learning

arXiv cs.CL·

Researchers train LLMs to compress long contexts before reasoning

arXiv cs.CL·
RefCon extracts agent memories without labels, scaling test-time compute · Modelwire