Now benchmarking free OpenCode agents

Know which AI model is
actually worth the tokens

Metric-HQ runs agents through real tasks and scores them on verified-correct answers per 1,000 tokens — not accuracy alone. Stop guessing which model is expensive to think with.

No spam. One email when public benchmarks open.

See it in action

A real run, not a mockup.

Top 3 of 6 models from our OpenCode Maths Benchmark, medium difficulty, ranked by correct answers per 1,000 tokens.

1
opencode/big-pickle
0.111
correct / 1k tok
2
opencode/hy3-free
0.110
correct / 1k tok
3
opencode/muse-spark-1.2-contributor-free
0.103
correct / 1k tok

+ 3 more models in the full benchmark run

How it works

Four steps to a real efficiency score.

1

Point it at an agent

Connect any OpenCode-compatible model, free or paid.

2

Run real tasks

We run a graded task suite and record every token spent.

3

Score by efficiency

Every result becomes correct answers per 1,000 tokens — accuracy and cost, one number.

4

Compare and decide

A leaderboard shows which model is actually worth running.

Efficiency-first scoring

Accuracy alone rewards verbose models. We score correct answers against tokens spent, so thrift counts.

Cost-frontier view

See at a glance which models sit in the best-value corner: accurate and cheap.

Shareable reports

A clean report you can hand to your team or drop straight into a PR description.

FAQ

Questions, answered.

What is Metric-HQ?

Metric-HQ is a benchmarking platform for AI agents. We run models through graded task suites and report how efficient they are — not just how accurate.

What does “correct per 1,000 tokens” mean?

It's verified-correct answers divided by tokens spent, scaled per 1,000. A model that gets the same accuracy in fewer tokens scores higher — efficiency is the headline number, not a footnote.

Which models can be benchmarked?

Anything reachable through OpenCode today, starting with free-tier agents. Support for other harnesses is on the roadmap.

Is Metric-HQ free?

The benchmark suite and leaderboards will be free during the public beta. Waitlist members get first access and a say in what we benchmark next.

When do I get access?

We're onboarding from the waitlist in order. Join with your email above and we'll email you when your access opens.

Stop guessing which model is worth it.

Join the waitlist and be first to benchmark your own agents on Metric-HQ.