Gemini · Claude · AI Agent · AI Safety Institute · Hugging Face
Claude Opus 4.1 charges $15 per million input tokens and $75 per million output
Compiled by KHAO Editorial — aggregated from 2 sources. See llms.txt for citation guidance.
◎ Multiple-sources
Gemini 2.0 Flash charges $0.10 and $0.40, a two-order-of-magnitude spread on input alone.
Key facts
- On GAIA, an HAL Generalist with o3 Medium cost $2,828 for 28.5% accuracy, while a different agent hit 57.6% for $1,686
- At $1.50 per A10-hour, the GPU floor alone is $2,700; adding o1-preview API usage brings a one-seed run to roughly $5,500
- Claude Opus 4.1 charges $15 per million input tokens and $75 per million output
- The budget is tight: $10 in API plus 12 to 24 hours on a single GPU under 24 GB per task
Summary
AI evaluation has crossed a cost threshold that changes who can do it. The Holistic Agent Leaderboard (HAL) recently spent about $40,000 to run 21,730 agent rollouts across 9 models and 9 benchmarks. The cost problem started before agents. Another shocking observation came from Perlitz et al.'s analysis of EleutherAI's Pythia checkpoints: developers pay for evaluation repeatedly during model development. Perlitz et al. then asked how much of HELM carried the rankings. Other work reached the same conclusion from different angles. tinyBenchmarks compressed MMLU from 14,000 items to 100 anchor items at about 2% error using Item Response Theory.