LLMs are stuck in a groupthink groove. This startup is trying to get them out.
Source published ·Modelwire updated
Original coverage: MIT Technology Review - AI ↗·How Modelwire adds context

The development
A startup is addressing a fundamental limitation in large language models: statistical clustering around predictable outputs. The piece demonstrates that major chatbots (Claude, ChatGPT, Gemini) exhibit measurable bias toward certain responses when asked for randomness, revealing how training data and sampling strategies create invisible guardrails. This groupthink problem affects downstream applications from creative generation to scientific simulation, where diversity of outputs matters. The startup's approach signals growing recognition that LLM behavior isn't truly stochastic but constrained by architectural and training choices that favor consensus outputs over genuine variance.
Modelwire’s AI-generated summary of coverage from MIT Technology Review - AI.
Modelwire analysis
Skeptical readOur AI-generated reading of the wider context and the next developments to watch.
The piece names the problem clearly but never discloses the startup's actual method, which makes it impossible to evaluate whether the fix is architectural, a post-processing layer, or a fine-tuning recipe. That omission is doing a lot of work here.
The groupthink finding lands differently when read alongside this week's coverage of Gemini. The Verge's piece on Google's smart speaker noted that Gemini's value depends on contextual reliability at scale, and if the model's outputs cluster toward consensus responses, that reliability is partly illusory. Separately, OpenAI's reported move toward three GPT-5.6 Pro variants (covered via The Decoder) suggests frontier labs are already segmenting models by behavior profile, which could be one commercial response to output homogeneity. Neither story directly validates this startup's approach, but together they show that output diversity is becoming a product-level concern, not just a research footnote.
If the startup publishes a reproducible benchmark showing measurable output variance improvement on a standard creative or simulation task within the next two quarters, that would distinguish a real technique from a positioning exercise. Absent that, this reads as a well-timed press cycle.
This interpretation is generated from the summary above and the archive coverage cited below. Our methodology · Report an error
Coverage behind this analysis
These archive entries ground the connection in our analysis. They are ordered by source publication date, with links to our coverage and the original sources.
·The Verge - AI
Google built a great smart speaker, but Gemini isn’t ready for it
Google's new smart speaker hardware represents a competitive response to Amazon's AI-powered Alexa refresh, but the device's value proposition hinges on Gemini's readiness for conversational, always-on interaction in the home. The gap between capable silicon and production-ready LLM integration exposes a recurring tension in consumer AI: hardware cycles move faster than model maturation. For the…
MentionsClaude · ChatGPT · Gemini · MIT Technology Review
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. The full content lives on technologyreview.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.