BenchMIRT: What are LLM benchmarks actually measuring?

Allen Institute for AI Version 1 original current

Imported from official source

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.

This version

Version
1 of 1
Recorded
September 20, 2026 19:52
Change
Initial
Content hash
4ef960848ee57721258a09c897a15038
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.