Llama · Hugging Face
Featuring Every Eval Ever Results on Hugging Face Model Pages
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
EEE launched in February 2026 as a project of the EvalEval Coalition, the first cross-institutional effort to improve how AI evaluation results get reported by both first and third party evaluators.
Key facts
- Since launching, the datastore on Hugging Face has grown to around 229,000 evaluation results across more than 22,000 models and 2,200 benchmarks, pulled from 31 different reporting formats
- Hugging Face launched Community Evals in February 2026 to decentralize how benchmark scores get reported on the Hub
- The same model on the same benchmark often returns different scores depending on who ran it and how; LLaMA 65B, for one, has been reported at both 63.7 and 48.8 on MMLU
- An Evaluation (MMLU-Pro) from EEE Datastore (a) cross-linked at the file level to a Hugging Face model card (b)
Summary
Evaluation results are how they measure model capabilities, compare models against each other, and reason about safety and governance, and yet they are scattered and hard to compare. The schema was built with feedback from researchers and policy researchers, and it takes in results from any source, so harness logs, leaderboard scrapes, and paper numbers all end up in the same shape. Now, it comes with better integration and attribution. This is new functionality for everyone who reports or reads evaluations, not only existing EEE contributors.