Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

A new Polish medical exam benchmark reveals that standard multiple-choice evaluation systematically inflates LLM medical competence. When researchers added 15,000 questions and structural modifications to reduce guessing artifacts, top performers like Qwen3.5-122B dropped 28-31 percentage points. This finding challenges the reliability of MCQA-based capability claims across the industry and suggests current benchmarking practices may mask significant gaps in clinical reasoning, forcing a reckoning over how medical AI systems should actually be validated.
Modelwire context
ExplainerThe 28-31 point drop isn't just a score correction, it's evidence that models may be exploiting answer-distribution artifacts and question phrasing patterns rather than demonstrating any underlying clinical knowledge. The benchmark's structural modifications, not just its scale, are doing the diagnostic work here.
This finding lands in direct conversation with the MedMisBench paper covered the same day ('Measuring Epistemic Resilience of LLMs Under Misleading Medical Context'), which showed models collapsing from 71% to 38% accuracy under adversarial context injection. Together, these two papers are building a consistent picture: high exam scores in medical AI are fragile in at least two distinct ways, once you stress the input structure and once you stress the reasoning context. Neither paper alone is decisive, but the convergence on the same day from independent research groups makes the pattern harder to dismiss as a single-lab artifact. The Polish benchmark adds a structural-bias angle that MedMisBench doesn't cover, and vice versa.
Watch whether benchmark maintainers for established medical evals like MedQA or USMLE-style suites adopt similar structural controls within the next two quarters. If they do and top-model scores hold, the Polish findings are partially scope-limited. If scores drop comparably, the problem is systemic.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsQwen3.5-122B · Polish medical exams benchmark · arXiv
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.