EpiBench tests LLM reasoning on antibody binding sites
Researchers have built EpiBench, a sequence-based benchmark that tests whether large language models can reason about epitopes, the molecular sites where antibodies bind antigens. This matters because epitope prediction directly influences antibody drug efficacy and resistance profiles, yet existing biomedical benchmarks skip this workflow entirely. The work exposes a capability gap: despite LLMs' proven strength in protein reasoning, their ability to infer binding sites from raw sequences alone remains unvalidated. For drug discovery teams and model developers, EpiBench establishes a new evaluation standard that bridges general protein understanding and real therapeutic constraints.
Modelwire context
ExplainerEpiBench isolates a specific failure mode: LLMs can reason about protein structure in abstract terms but struggle to infer the precise molecular binding sites that determine antibody efficacy. This is not a general protein understanding problem but a localization problem that existing benchmarks don't measure.
This follows the pattern established by onepot-Bench 0 (August 3rd) and MedUPS (August 2nd), which both identified gaps between what LLMs claim to do and what they actually do in domain-specific workflows. Like onepot-Bench's focus on wet-lab chemistry judgment and MedUPS's emphasis on sequential clinical reasoning, EpiBench targets the intermediate step that matters most in practice rather than the final output. The difference here is that epitope prediction is a pure sequence inference task with no temporal component, making it a cleaner test of whether LLMs have learned the underlying binding physics or are pattern-matching on training data.
If EpiBench results correlate with actual antibody success rates in published clinical trials (comparing predictions on approved drugs vs. failed candidates), that validates the benchmark as predictive. If performance stays flat across model scales above 7B parameters, that suggests the gap is architectural rather than just data-driven, which would matter for how drug teams should weight LLM assistance in their pipelines.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.