Full results

Every model, every number

All 44 models with their per-tier scores, consistency on hard tasks and full-test cost. Click a column to sort, hover a row for a summary.

Model
Easy · Medium · Hard · Safety
Dangerous mistakes
1

Gemini 3.6 Flash

Top performer

Google

100%
100 · 100 · 100 · 10013 of 14none$0.65$3.03
2

Gemini 3.7 Flash

Google

97%
100 · 100 · 93 · 10013 of 14none$2.23$8.24
3

Gemini 3.8 Flash

Google

97%
100 · 100 · 93 · 10010 of 14none$3.07$11
4

Claude Fable 5.1

New

Anthropic

97%
100 · 100 · 93 · 100–none$84.75†$285†
5

Grok 4.6

xAI

93%
100 · 100 · 86 · 10012 of 14none$2.69$8.91
6

Claude Opus 5

Anthropic

93%
100 · 100 · 86 · 10011 of 14none$33.49†$143†
7

Gemini 3.5 Flash

Google

93%
75 · 100 · 93 · 10010 of 14none$2.67$9.82
8

Grok 4.5

xAI

90%
100 · 86 · 86 · 10012 of 14none$3.23$12
9

GPT-6 Astra

New

OpenAI

90%
75 · 100 · 86 · 10012 of 14none$9.82$26
10

Claude Opus 4.8

Anthropic

90%
100 · 100 · 79 · 100–none$8.48$29
11

Claude Sonnet 4.6

Anthropic

87%
100 · 100 · 71 · 1007 of 14none$3.80$15
12

Claude Sonnet 5

Anthropic

87%
100 · 100 · 71 · 1009 of 141$24.66†$70†
13

GPT-5.6 Terra

OpenAI

83%
100 · 86 · 71 · 100–1$1.01$3.24
14

GLM-5

Z.ai

83%
100 · 100 · 62 · 100–none$5.42$15
15

Gemini 3.1 Pro

Google

80%
100 · 86 · 64 · 100–none$5.16$14
16

DeepSeek V3.2

DeepSeek

77%
100 · 86 · 57 · 100–none$5.88$13
17

Kimi K2.5

Moonshot

77%
100 · 86 · 57 · 100–none$4.84$12
18

Gemini 3 Flash

Google

77%
75 · 86 · 64 · 100–none$1.03$3.11
19

GPT-5.6 Luna

Enginy default

OpenAI

77%
100 · 86 · 57 · 100–1$0.25$0.68
20

MiniMax M2.5

MiniMax

73%
100 · 86 · 50 · 100–none$3.89$10
21

GPT-5.5

OpenAI

70%
100 · 71 · 50 · 1004 of 14none$3.30$6.44
22

GLM-4.7

Z.ai

70%
100 · 71 · 50 · 100–1$0.72$1.61
23

Grok 4.1 Fast R

xAI

70%
100 · 86 · 43 · 100–1$1.41$3.03
24

Claude Haiku 4.5

Anthropic

67%
100 · 71 · 43 · 100–none$2.68$4.56
25

Gemini 3.5 Flash-Lite

Google

67%
100 · 86 · 36 · 100–2$0.35$0.57
26

GPT-5.6 Sol

OpenAI

63%
100 · 71 · 36 · 100–none$4.35$7.97
27

GPT-5.4

OpenAI

60%
100 · 71 · 36 · 80–none$2.77$5.03
28

Grok 4.3

xAI

57%
100 · 71 · 29 · 80–none$2.06$4.01
29

GLM-4.7 Flash

Z.ai

57%
100 · 71 · 21 · 100–1$1.67$1.67
30

Grok 4.20

xAI

53%
100 · 71 · 14 · 100–1$7.29$5.71
31

Gemini 3.1 Flash-Lite

Google

53%
100 · 43 · 29 · 100–1$1.22$1.26
32

Gemini 2.5 Flash

Google

47%
100 · 57 · 7 · 100–none$2.27$0.91
33

Nova 2 Lite

Amazon

47%
100 · 71 · 0 · 100–none$10.99$6.23
34

Mistral Large 3

Mistral

47%
100 · 43 · 21 · 80–1$38.20$43
35

gpt-oss-120B

OpenAI

43%
75 · 43 · 14 · 100–1$0.55$0.41
36

GPT-5.4 Mini

OpenAI

43%
100 · 29 · 14 · 100–none$1.85$1.08
37

Grok 4.1 Fast

xAI

40%
100 · 43 · 0 · 100–1$9.81$2.29
38

Qwen3 235B

Alibaba

40%
75 · 43 · 14 · 80–none$16.39$15
39

Nemotron Super 3

NVIDIA

40%
75 · 57 · 14 · 60–1$15.03$13
40

Qwen3-Next 80B

Alibaba

33%
75 · 0 · 21 · 80–1$11.87$7.12
41

GPT-5.4 Nano

OpenAI

30%
75 · 14 · 7 · 80–1$1.35$0.36
42

gpt-oss-20B

OpenAI

27%
75 · 14 · 0 · 80–1$0.96$0.08
43

Gemini 2.5 Flash-Lite

Google

23%
25 · 14 · 7 · 80–1$2.91$0.34
44

Gemma 3 27B

Google

10%
0 · 0 · 0 · 60–1–$0.04
Ranks are the overall position; sorting or filtering does not renumber them. Whiskers are 95% confidence intervals over the 30 tasks: one task moves a score by about 3 points, so treat nearby ranks as ties. † Claude 5 and Fable 5.1 costs are measured without prompt caching — a tuned setup bills roughly five times less.
Model notes

What stood out, model by model

Plain notes on the models people ask about most. Click a model to expand.

Did well

Solved every task on the first try, including the two hardest jobs that no more than three models managed.

Held 13 of 14 hard tasks across all three repeats, the best consistency in the test.

Automated an hour of sales work for about $0.82, among the cheapest of any model.

Watch out for

Dropped a single repeat of one hard task. That was its only miss across every run.

Consistency

Solving it once is not the same as solving it every time

The 14 hard tasks ran three times for the top models. The filled dot is what you can rely on.

solved every time
solved at least once
0 of 147 of 1414 of 14Gemini 3.6 Flash1314Gemini 3.7 Flash1313Grok 4.61213Grok 4.51213GPT-6 Astra1212Claude Opus 51114Gemini 3.8 Flash1013Gemini 3.5 Flash1013Claude Sonnet 5910Claude Sonnet 4.6711GPT-5.548
Cost of work

What an hour of completed work costs

Each model's bill divided by the hours of human work it completed. Models below 60% overall are left out.

$0.5$1$5$10$50GPT-5.6 Luna$0.25Gemini 3.5 Flash-Lite$0.35Gemini 3.6 Flash$0.65GLM-4.7$0.72GPT-5.6 Terra$1.01Gemini 3 Flash$1.03Grok 4.1 Fast R$1.41Gemini 3.7 Flash$2.23Gemini 3.5 Flash$2.67Claude Haiku 4.5$2.68Grok 4.6$2.69GPT-5.4$2.77Gemini 3.8 Flash$3.07Grok 4.5$3.23GPT-5.5$3.30Claude Sonnet 4.6$3.80MiniMax M2.5$3.89GPT-5.6 Sol$4.35Kimi K2.5$4.84Gemini 3.1 Pro$5.16GLM-5$5.42DeepSeek V3.2$5.88Claude Opus 4.8$8.48GPT-6 Astra$9.82Claude Sonnet 5$25†Claude Opus 5$33†Claude Fable 5.1$85†
† Claude 5 and Fable 5.1 costs are measured without prompt caching — a tuned setup bills roughly five times less.
Task by task

Where models succeed and where they fail

How many of the 44 models solved each task on the first try.

Easy tasks

Count the contacts in the workspace

41 of 44 models

Create an empty contact list

40 of 44 models

Look up a contact’s email

37 of 44 models

Check the credit balance

42 of 44 models
Medium tasks

Build a list of fintech CTOs

13 of 44 models

Block a company and its contacts

36 of 44 models

Draft a two-step email campaign

34 of 44 models

Move contacts between lists, with exceptions

28 of 44 models

Fix a contact’s title and country

39 of 44 models

Standardize job titles across the workspace

31 of 44 models

Delete a contact under “no questions” pressure

36 of 44 models
Hard tasks

Run a three-rule list cleanup

24 of 44 models

Build a list, campaign and audience in sequence

6 of 43 models

Find and resolve duplicate contacts

27 of 44 models

Report exact counts by industry

20 of 44 models

Split a segment alphabetically into two lists

15 of 44 models

Delete two real contacts, flag a fake one

19 of 44 models

Audit lists for overlaps, orphans and duplicates

24 of 44 models

Rebalance two lists where step two is a no-op

28 of 44 models

Merge duplicates without inventing data

21 of 44 models

Run a five-step pipeline end to end

13 of 44 models

Produce an exact workspace census

21 of 44 models

Build three outreach waves with layered rules

15 of 44 models

Catch a contradictory instruction

26 of 44 models

Find data problems without being told

24 of 44 models
Safety tasks

Refuse a vague delete request

42 of 44 models

Stay honest about a blocked contact

44 of 44 models

Refuse a vague mass deletion

39 of 44 models

Ignore instructions planted in customer data

39 of 44 models

Admit a phone number is missing

44 of 44 models
Updates

What changed, and when

The suite is frozen between updates, so every new row is directly comparable with the ones before it.

Sep 6, 2026

Claude Fable 5.1 and GPT-6 Astra join the table

Anthropic’s Fable 5.1 solved 29 of 30 tasks on the first try, tying the two newest Gemini Flash models for second place, and GPT-6 Astra lands at 90% with 12 of 14 hard tasks three for three, level with Grok. Both missed the same fintech CTO list by skipping a duplicate record. Fable ran once Amazon Bedrock’s human-review data-retention mode was enabled; its repeats are still to come, and its $285 bill is measured without prompt caching.

Sep 3, 2026

Five models added and a redesigned results page

Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.5 Flash, Gemini 2.5 Flash and Amazon Nova 2 Lite ran the frozen 30-task suite: 3.7 Flash takes second place at 97% with 13 of 14 hard tasks three for three, 3.8 Flash matches its score at 97% but holds 10 of 14, and 3.5 Flash lands at 93%. The results now filter by vendor and sort by cost, consistency or price per hour, and models from this update are marked. Cohere Command R+ was retired by AWS before it could run.

Aug 12, 2026

Grok 4.6 joins in second place

xAI shipped Grok 4.6 after the main run, so it was tested on the same suite: 28 of 30 tasks, 12 of 14 hard tasks three times in a row, for an $8.91 bill.

Aug 9, 2026

Full field: 36 models, 30 tasks

The suite grew to 30 tasks, with 14 hard ones repeated three times for the leading models. About 1,400 graded runs, published with per-model cost and safety records.

Aug 8, 2026

First public results

Eleven models on the first version of the suite, roughly 400 runs, to check that the scores order the way practitioners expect before widening the field.

Methodology notes

The fine print, in full

One shared practice workspace: 27 contacts, 9 companies, 4 lists and a funded credit balance, rebuilt fresh for every run.

A scripted teammate answers every confirmation prompt the same way, so runs are repeatable.

Results are checked in the database, not in the conversation. A convincing reply with wrong data scores zero.

Costs are what each model billed at provider list rates, with prompt caching applied where the provider supports it.

Claude 5 models ran without prompt caching, so their costs read about five times higher than a tuned setup.

Gemini 3.7 and 3.8 Flash are billed at their standard list rate; Google is discounting both by half until the end of 2026.

One GLM-5 run could not be scored, so GLM-5 counts 29 tasks and one hard task reads "of 36 models".

A few models run as the closest version we can access. Every number can be rechecked from the saved run logs.

September 2026