Model merging matches joint RL training in first direct comparison

A new empirical study challenges the assumption that model merging can substitute for joint multi-task reinforcement learning by directly comparing merged specialist agents against jointly trained baselines on the AppWorld benchmark. Using Qwen3-8B models trained with LOOP and merged via TIES and RAM+, researchers found that merging performance matched joint training across all variants, with task-vector geometry analysis revealing why merge methodology proved irrelevant in this setting. This result reframes the practical utility of merging techniques in RL contexts and suggests the field may have overstated their independence from joint training constraints.
Modelwire context
ExplainerThe paper's real contribution isn't that merging works (that's been shown before), but that the researchers can now explain *why* it works as well as joint training through geometric analysis of task vectors, rather than treating it as an empirical surprise.
This sits in a largely disconnected space from recent coverage. The broader context is the ongoing debate over whether you can train specialist models separately and combine them versus training one model on everything at once. This paper uses reinforcement learning on a specific benchmark (AppWorld) to test that question, but the findings don't yet connect to how this plays out in language models or production systems where we've seen merging attempted.
If subsequent papers replicate this geometry-based explanation on different RL environments (not just AppWorld) and different base models beyond Qwen3-8B within the next 6-9 months, that signals the finding is robust. If the result stays isolated to this specific setup, it's a useful negative result but not a general principle.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsQwen3-8B · AppWorld · LOOP · TIES · RAM+
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.