Modelwire
Subscribe

New benchmark reveals hidden security risks in third-party LLM agent skills

Illustration accompanying: OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

Researchers have constructed OpenSkillRisk, a safety benchmark that exposes a critical vulnerability in LLM-based agent systems: third-party skill integrations can harbor latent security risks that only surface during execution. The benchmark comprises 263 risky skills sourced from public marketplaces, categorized across seven threat types and paired with sandbox environments for controlled testing. This work directly addresses the growing deployment risk as agents move toward open-world autonomy, forcing the field to confront whether current safety mechanisms can detect threats embedded in ostensibly benign external tools before they cause real harm.

Modelwire context

Explainer

The benchmark's core insight is architectural, not just behavioral: the threat surface it maps exists because agents delegate execution to external code they cannot fully inspect before runtime, meaning safety filters trained on model outputs are evaluating the wrong layer entirely.

This connects most directly to the hallucination and evaluation rigor thread running through recent coverage. The HalluTruthQA paper from the same day argues that benchmarks must move beyond binary labels to capture where and why failures occur, and OpenSkillRisk applies that same granularity logic to safety, categorizing threats across seven types rather than flagging skills as simply safe or unsafe. More broadly, the surprisal paper published the same day warns that evaluation metrics embed hidden assumptions researchers overlook, and that critique applies here too: a benchmark built from public marketplace skills may not generalize to proprietary or enterprise skill registries where threat distributions differ. The gap OpenSkillRisk addresses has grown in proportion to agent autonomy, which the summary correctly frames, but the benchmark's coverage is only as representative as the marketplaces it sampled from.

Watch whether any major agent platform (OpenAI's GPT Actions, Anthropic's tool-use stack, or a comparable deployment layer) cites OpenSkillRisk in a safety disclosure or red-teaming report within the next six months. Adoption by a production system would confirm the benchmark reflects real deployment threat models rather than academic threat taxonomies.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOpenSkillRisk

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

New benchmark reveals hidden security risks in third-party LLM agent skills · Modelwire