SenFlow: Inter-Sentence Flow Modeling for AI-Generated Text Detection in Hybrid Documents

Researchers have reframed sentence-level AI-detection in mixed-authorship documents as a structured prediction problem, moving beyond isolated classification to capture how generated and human text interact within a document. The new MOSAIC benchmark tests detection against recent frontier models (DeepSeek-V3.2, Kimi K2) using stricter quality controls than prior datasets. SenFlow, which models inter-sentence dependencies through graph propagation and CRF decoding, achieves state-of-the-art results. This work matters because hybrid documents are becoming the norm in research and publishing, and detection systems that ignore sequential context will fail as LLMs improve at local coherence.
Modelwire context
ExplainerThe deeper contribution is the MOSAIC benchmark itself: by specifically including outputs from frontier models like DeepSeek-V3.2 and Kimi K2 with stricter quality controls, the researchers are directly addressing the benchmark staleness problem that has quietly undermined prior detection work, where training and evaluation sets predate the models now in widespread use.
This connects directly to the piece on 'Which Sections of a Research Paper Best Reveal Its Research Methods,' which also treats academic documents as structured objects where position and sequence carry meaning beyond raw content. Both papers are pushing against the same flat-text assumption that has dominated NLP tooling for academic publishing. More broadly, the detection framing here complements the concern raised in 'Written by AI, Managed by AI' about how AI-generated content accumulates and drifts in long-horizon workflows, since hybrid documents are precisely the output that kind of workflow produces. SenFlow is, in a sense, trying to read the seams those workflows leave behind.
Watch whether MOSAIC gets adopted as an evaluation target by the authorship verification community within the next two conference cycles. If it does not, the benchmark risks the same obsolescence it was designed to correct, as frontier models continue to close the local coherence gap SenFlow currently exploits.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSenFlow · MOSAIC · DeepSeek-V3.2 · Kimi K2 · PubMed · XSum
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.