Modelwire
Subscribe

Upstage scales Solar Open to 250B parameters with 1M-token context

Illustration accompanying: Solar Open 2 Technical Report

Upstage has scaled Solar Open to 250 billion parameters using a mixture-of-experts architecture designed for long-context agent reasoning. The model achieves a 1M-token context window through a hybrid attention mechanism combining softmax and linear layers, enabling entire agent trajectories to fit in memory. By initializing from Solar Open 1 and leveraging higher-quality training data, the team maintained efficiency under fixed compute budgets. This represents a significant step toward production-grade models capable of sustained multi-step reasoning, directly addressing a key bottleneck in agentic AI systems.

Modelwire context

Analyst take

The 1M-token context window is the headline, but the more consequential detail is the architectural choice to combine softmax and linear attention layers rather than adopt pure linear attention, a tradeoff that preserves quality on short contexts while extending range, and one that carries real inference cost implications Upstage has not yet publicly benchmarked at production throughput.

The attention efficiency problem Solar Open 2 is navigating sits directly alongside what ELSAA addressed earlier this week: the quadratic cost of attention at long sequence lengths is the binding constraint, and hybrid approximation is emerging as the practical answer rather than any single clean solution. Meanwhile, the MoE scaling choice puts Upstage in the same infrastructure territory as the SLAI T-Rex work on DeepSeek-V4, where trillion-parameter MoE training economics are being stress-tested across different hardware stacks. What remains unaddressed is the safety surface that OpenSkillRisk flagged: a 250B-parameter model with 1M-token context designed explicitly for agentic trajectories is precisely the deployment profile where third-party skill risks compound.

Watch whether Upstage publishes inference throughput numbers at full 1M-token context on commodity hardware within the next 60 days. If they do not, the context window claim is a training-time capability, not a deployable one.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsUpstage · Solar Open 2 · Solar Open 1 · Mixture-of-Experts

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Solar Open 2 Technical Report”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Upstage scales Solar Open to 250B parameters with 1M-token context · Modelwire