All benchmarks in the reference index — sorted, sourced, searchable. Reference index; measurement, not certification.
| Name | Region | Binding | Status | Contam. |
|---|---|---|---|---|
| AgentBench | Global | no | active | low |
| AGIEval | Global | no | active | medium |
| AILuminate (MLCommons AI Safety Benchmark) | Global | no | active | medium |
| ARC — Abstraction and Reasoning Corpus | Global | no | active | medium |
| BIG-bench — Beyond the Imitation Game | Global | no | active | medium |
| BoolQ | Global | no | active | medium |
| Chatbot Arena (LMSYS) | Global | no | active | low |
| CommonsenseQA | Global | no | active | medium |
| GAIA — General AI Assistants | Global | no | active | low |
| GLUE — General Language Understanding Evaluation | Global | no | active (superseded by SuperGLUE) | high |
| GPQA — Graduate-Level Google-Proof Q&A | Global | no | active | low |
| GSM8K — Grade School Math | Global | no | active | high |
| HarmBench | Global | no | active | low |
| HellaSwag | Global | no | active | medium |
| HELM — Holistic Evaluation of Language Models | Global | no | active | medium |
| HumanEval | Global | no | active | high |
| LAMBADA | Global | no | active | medium |
| LiveBench | Global | no | active | low |
| MATH (Hendrycks) | Global | no | active | medium |
| MathVista | Global | no | active | low |
| MMLU — Massive Multitask Language Understanding | Global | no | active | high |
| MMLU-Pro | Global | no | active | low |
| MMMU — Massive Multi-discipline Multimodal Understanding | Global | no | active | low |
| OpenBookQA | Global | no | active | medium |
| PIQA — Physical Interaction QA | Global | no | active | medium |
| RACE | Global | no | active | medium |
| SciQ | Global | no | active | medium |
| SQuAD 2.0 | Global | no | active | medium |
| SuperGLUE | Global | no | active | medium |
| SWE-bench | Global | no | active | high |
| ToolBench | Global | no | active | low |
| TruthfulQA | Global | no | active | medium |
| VQA v2.0 | Global | no | active | medium |
| WebArena | Global | no | active | low |
| WinoGrande | Global | no | active | medium |
| WMDP — Weapons of Mass Destruction Proxy | Global | no | active | low |
Honest register: reference index — describes and sources; it does not certify or score trust. Measurement, not certification. Data via MCP (drum_search/drum_get/drum_crosswalk) + feeds/reg_events.json.
EAT 7-box mission: PARTIAL — 4/7 true, 1 partial, 2 false (honest register). true=boarded,ci,measured,mirrored · partial=chained · false=anchored,signed. Measurement, not certification.