
New benchmark tackles the missing half of factuality in language models
Researchers introduce a two-level meta-rubric framework to address a critical gap in LLM evaluation: measuring factual completeness alongside precision. Current benchmarks excel at catching false claims but fail to assess whether generated responses omit important information. GAMUT tackles this by modeling facts as hierarchical structures with dependencies and open-ended coverage requirements, moving beyond simple boolean checklists. This work matters because production systems need both correctness and comprehensiveness metrics, and the framework's approach to structured fact relationships could reshape how teams validate long-form generation quality across summarization, QA, and reasoning tasks.62


























.jpg)

