Keyword benchmarks mask tool-use failures in small language models
Researchers expose a critical flaw in how tool-use capabilities are measured in language models: keyword-matching benchmarks systematically overstate performance, especially in smaller models. By comparing two Spanish security models with identical architecture but different training regimes, they show that lenient metrics mask complete failure on actual tool invocation. The work introduces a diagnostic hierarchy using verbatim reproduction checks and token-level probes to separate genuine tool use from spurious pattern matching. This matters because inflated tool-use claims have become common in model evaluation, and the proposed cheap diagnostics offer a practical path toward more honest capability assessment across the industry.
Modelwire context
Skeptical readThe paper frames keyword-harness failure as a measurement problem, but doesn't establish whether the models were ever claimed to have genuine tool-use capability in the first place. The diagnostic hierarchy is presented as novel, yet the core insight (small models memorize patterns without executing) is not new to practitioners building security tooling.
This connects directly to KaliBench (released the same day), which also targets tool-use precision in security contexts but takes the opposite approach: instead of diagnosing why benchmarks fail, KaliBench builds a fine-grained dataset with verifiable rewards tied to actual command execution. Where this paper says 'current metrics are too lenient,' KaliBench says 'we need executable ground truth.' The tension matters: if keyword-matching was always a proxy metric, why did the field treat it as a capability measure rather than a signal? The earlier work on CoT-Pass@k and Mathematical Primitives both expose similar gaps between surface-level correctness and actual reasoning, suggesting the real problem is that the community conflates metric convenience with capability validation.
If the authors apply their diagnostic hierarchy to the KaliBench dataset and find that models passing their token-level probes still fail on actual Kali command generation, that confirms the diagnostics are real. If instead the two benchmarks show high correlation, the paper's claim that current metrics 'systematically overstate' performance collapses into a tautology about measurement granularity rather than a flaw in the field's evaluation practice.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSpanish security language models · 661.6M parameter model · 1,109M parameter model
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.