Modelwire
Subscribe

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

Illustration accompanying: Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

Researchers analyzed 14,727 security and privacy queries from a real-world dataset of 3.2M LLM conversations, establishing the first empirical baseline for how users actually seek S&P guidance from AI systems rather than relying on expert-constructed test cases. This work exposes a critical gap in LLM safety evaluation: current benchmarks miss the messy, context-dependent questions users genuinely ask, meaning response quality assessments may not reflect real-world harm or helpfulness. The categorization framework surfaces which S&P domains users struggle with most, informing both model training priorities and the design of safer guardrails for high-stakes advice.

Modelwire context

Explainer

The study's most underappreciated contribution is methodological: by drawing from WildChat's 3.2 million real conversations rather than curated test suites, it reveals that safety evaluations have been graded on a curve, measuring model performance against questions experts thought to ask rather than the ones users actually bring.

This connects directly to the Handlebars templating vulnerability covered in 'Structural Role Injection in Handlebars-Templated LLM Prompts' from the same day. That paper exposed a gap between documented security assumptions and real attack surface in production systems. This paper exposes the same structural problem one layer up: the evaluation frameworks meant to catch dangerous model behavior are themselves calibrated against idealized inputs, not the messy, context-dependent queries that arrive in production. Both papers, arriving together, suggest the field has a recurring habit of testing for the threats it anticipated rather than the ones that actually materialize.

Watch whether any major model provider cites this dataset when updating their safety benchmarks in the next two quarters. If the WildChat-derived categories get absorbed into standard red-teaming protocols, that confirms the field treated this as a genuine calibration problem rather than an academic footnote.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsWildChat · Large Language Models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond · Modelwire