
LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards
Researchers introduce LongTraceRL, a reinforcement learning framework that tackles a persistent weakness in LLMs: extracting and reasoning over relevant information buried in lengthy documents. The method improves on prior RLVR approaches by constructing high-fidelity distractors from search agent behavior and replacing sparse outcome rewards with fine-grained rubric-based signals that supervise intermediate reasoning steps. This addresses a real bottleneck in production retrieval-augmented systems, where models struggle to distinguish signal from noise across long contexts, making the work relevant to anyone building search or QA infrastructure at scale.62
























