Copilot · Claude Code · Claude · GitHub · Codex · GPT · GitHub Blog
Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
While the model provides the raw intelligence, the harness shapes how effectively that intelligence is applied.
Key facts
- Below they report their latest results for a subset of the benchmarks they track, across four leading models: Claude Sonnet 4.6, Claude Opus 4.7, GPT‑5.4, and GPT‑5.5
- Throughout, they compare GitHub Copilot CLI against the model-vendor harnesses that ship those models natively: Claude Code for Sonnet 4.6 and Opus 4.7, and Codex CLI for GPT‑5.4 and GPT‑5.5
- The GitHub Copilot agentic harness supports 20+ frontier models across the GPT, Claude, Gemini, and MAI families, plus bring your own key for open‑source and local models
- Every marker is one agent-and-model configuration on TerminalBench 2.0, with resolution rate on the vertical axis and dollar cost per task on the horizontal axis
Summary
The tools, context, and workflow are orchestrated by the harness. In this post, they'll present data showing the efficiency and performance of the GitHub Copilot agentic harness across a wide range of agentic software engineering tasks. The team continuously evaluate the capability and efficiency of the GitHub Copilot agentic harness through a combination of public and internally developed benchmarks. The team control as many variables as possible to evaluate the performance of GitHub Copilot’s harness compared to the model provider’s harness: use the same model, the same benchmark task, normalized on context window, reasoning efforts, tool selection, and MCP servers.