Self-Stigma Is Not a Monolith, but Generic Empathy Is: Persona-Conditioned LLM Support for People Who Use Drugs

Researchers developed a persona-aware LLM system that tailors support conversations to distinct psychological profiles of people struggling with substance use, moving beyond one-size-fits-all empathy. Using latent profile analysis on Reddit data, they identified four self-stigma personas and built classifiers (macro-F1 0.74) that outperform few-shot LLM baselines at persona detection from limited posting history. Clinical expert evaluation across three LLMs exposed a critical gap: persona-matched responses diverged from what experts deemed appropriate, surfacing a fundamental challenge in aligning LLM behavior to nuanced clinical contexts where generic responses may actively harm vulnerable populations.
Modelwire context
ExplainerThe headline finding isn't that persona detection works reasonably well (macro-F1 0.74 is solid but not remarkable). It's that even when the system correctly identifies a user's psychological profile, the LLM's persona-matched response still diverges from what clinical experts consider appropriate, meaning the bottleneck isn't classification accuracy but something deeper about how LLMs model therapeutic appropriateness.
This connects directly to the 'Do LLM Embedding Spaces Recover Expert Structure?' paper covered the same day, which found that LLM representations diverge from expert-defined relationships even when surface-level separability looks fine. Both papers are essentially reporting the same structural problem from different angles: high task performance on a proxy metric does not mean the model has internalized the domain knowledge practitioners actually need. Together they suggest a pattern worth naming, which is that validation methodology in clinical NLP is systematically optimizing for the wrong signals.
The critical next test is whether any of the three evaluated LLMs can close the expert-alignment gap when given explicit clinical guidelines as system prompts rather than persona labels alone. If structured clinical framing doesn't move expert ratings meaningfully, that would indicate the problem is in the model's underlying representations, not just prompt design.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsLatent Profile Analysis · Bayesian classifiers · recurrent neural networks · Reddit · LLM
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.