SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection

Researchers have identified a phase-transition pattern in how large language models fail at stance detection as task complexity increases. The SICI framework, a seven-dimensional diagnostic tool, reveals that model errors follow distinct regimes: simple cases trigger over-prediction of opposition, mid-range examples destabilize, and hard cases collapse toward neutral predictions. This finding matters because it suggests current prompting techniques, retrieval augmentation, and reasoning chains cannot overcome fundamental complexity thresholds. The pattern holds across GPT-3.5 and GPT-4o-mini, indicating the limitation is structural rather than model-specific, and points toward the need for architectural or training-level solutions rather than prompt engineering.
Modelwire context
ExplainerThe more pointed finding buried in the SICI paper is that chain-of-thought prompting and retrieval augmentation, two of the most commonly deployed fixes for LLM reasoning failures, were explicitly tested and failed to overcome the complexity thresholds. That rules out a large class of practitioner workarounds before they're tried.
This connects directly to a pattern emerging across several recent papers on this site. The reward model piece ('Understanding helpfulness and harmless tension') found that competing objectives hit hard architectural limits at the neuron level, not the prompting level. SICI is making a structurally similar argument from a different angle: that certain failure modes are baked into the model rather than addressable at inference time. The Arabic-Hebrew cognate work ('When Similar Means Different') adds a third data point, showing models collapse toward surface heuristics under semantic pressure. Taken together, these papers are converging on a shared claim that current architectures have load-bearing weaknesses that scaling and prompting cannot paper over.
The key test is whether the SICI regime-shift pattern replicates on models beyond the GPT-3.5 and GPT-4o-mini family, particularly on open-weight models where architectural internals are inspectable. If the same three-regime collapse appears in Llama or Mistral variants, the structural explanation gains real traction.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT-3.5 · GPT-4o-mini · SemEval-2016 · VAST · SICI
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.