Modelwire
Subscribe

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Illustration accompanying: From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Researchers have mapped how essay quality emerges within the hidden layers of large language models, revealing that scoring signals are encoded in linearly separable form across eight different LLMs. By probing representations across three datasets spanning English and Portuguese, the work demonstrates that quality judgments form progressively through model depth, persist across different prompting approaches, and partially generalize across essay types despite varying rubrics. This interpretability finding matters because it suggests LLM-based scoring systems operate on learnable, transferable quality concepts rather than brittle pattern matching, potentially improving reliability in automated assessment systems used at scale.

Modelwire context

Explainer

The key finding isn't just that LLMs can score essays, it's that quality representations form gradually across model depth in a geometrically structured way, which means probing these layers could eventually let practitioners audit or correct scoring behavior without retraining the full model. That diagnostic potential is what the summary undersells.

This connects directly to the same-day arXiv paper on psychological profiles of LLMs ('Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact'), which found that apparent model traits reflect measurement artifacts rather than stable internal states. That paper undermines behavioral evaluation; this essay-scoring paper pushes in the opposite direction, arguing that at least some internal representations are stable and meaningful. Together they sketch a more complicated picture: not all LLM internals are noise, but distinguishing signal from artifact requires the kind of layer-level probing methodology this paper demonstrates. The tension between those two findings is worth holding onto as interpretability research matures.

Watch whether any of the three datasets used here (ASAP++, CSEE, ENEM) get adopted as benchmarks in future automated essay scoring audits. If independent replication confirms cross-lingual generalization between the English and Portuguese corpora, the transferability claim becomes a foundation for multilingual scoring deployment rather than a footnote.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsASAP++ · CSEE · ENEM · LLMs

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models · Modelwire