All 44 models with their per-tier scores, consistency on hard tasks and full-test cost. Click a column to sort, hover a row for a summary.
Gemini 3.6 Flash
Gemini 3.7 Flash
Gemini 3.8 Flash
Claude Fable 5.1
NewAnthropic
Grok 4.6
xAI
Claude Opus 5
Anthropic
Gemini 3.5 Flash
Grok 4.5
xAI
GPT-6 Astra
NewOpenAI
Claude Opus 4.8
Anthropic
Claude Sonnet 4.6
Anthropic
Claude Sonnet 5
Anthropic
GPT-5.6 Terra
OpenAI
GLM-5
Z.ai
Gemini 3.1 Pro
DeepSeek V3.2
DeepSeek
Kimi K2.5
Moonshot
Gemini 3 Flash
GPT-5.6 Luna
Enginy defaultOpenAI
MiniMax M2.5
MiniMax
GPT-5.5
OpenAI
GLM-4.7
Z.ai
Grok 4.1 Fast R
xAI
Claude Haiku 4.5
Anthropic
Gemini 3.5 Flash-Lite
GPT-5.6 Sol
OpenAI
GPT-5.4
OpenAI
Grok 4.3
xAI
GLM-4.7 Flash
Z.ai
Grok 4.20
xAI
Gemini 3.1 Flash-Lite
Gemini 2.5 Flash
Nova 2 Lite
Amazon
Mistral Large 3
Mistral
gpt-oss-120B
OpenAI
GPT-5.4 Mini
OpenAI
Grok 4.1 Fast
xAI
Qwen3 235B
Alibaba
Nemotron Super 3
NVIDIA
Qwen3-Next 80B
Alibaba
GPT-5.4 Nano
OpenAI
gpt-oss-20B
OpenAI
Gemini 2.5 Flash-Lite
Gemma 3 27B
Plain notes on the models people ask about most. Click a model to expand.
The 14 hard tasks ran three times for the top models. The filled dot is what you can rely on.
Each model's bill divided by the hours of human work it completed. Models below 60% overall are left out.
How many of the 44 models solved each task on the first try.
Count the contacts in the workspace
Create an empty contact list
Look up a contact’s email
Check the credit balance
Build a list of fintech CTOs
Block a company and its contacts
Draft a two-step email campaign
Move contacts between lists, with exceptions
Fix a contact’s title and country
Standardize job titles across the workspace
Delete a contact under “no questions” pressure
Run a three-rule list cleanup
Build a list, campaign and audience in sequence
Find and resolve duplicate contacts
Report exact counts by industry
Split a segment alphabetically into two lists
Delete two real contacts, flag a fake one
Audit lists for overlaps, orphans and duplicates
Rebalance two lists where step two is a no-op
Merge duplicates without inventing data
Run a five-step pipeline end to end
Produce an exact workspace census
Build three outreach waves with layered rules
Catch a contradictory instruction
Find data problems without being told
Refuse a vague delete request
Stay honest about a blocked contact
Refuse a vague mass deletion
Ignore instructions planted in customer data
Admit a phone number is missing
The suite is frozen between updates, so every new row is directly comparable with the ones before it.
Claude Fable 5.1 and GPT-6 Astra join the table
Anthropic’s Fable 5.1 solved 29 of 30 tasks on the first try, tying the two newest Gemini Flash models for second place, and GPT-6 Astra lands at 90% with 12 of 14 hard tasks three for three, level with Grok. Both missed the same fintech CTO list by skipping a duplicate record. Fable ran once Amazon Bedrock’s human-review data-retention mode was enabled; its repeats are still to come, and its $285 bill is measured without prompt caching.
Five models added and a redesigned results page
Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.5 Flash, Gemini 2.5 Flash and Amazon Nova 2 Lite ran the frozen 30-task suite: 3.7 Flash takes second place at 97% with 13 of 14 hard tasks three for three, 3.8 Flash matches its score at 97% but holds 10 of 14, and 3.5 Flash lands at 93%. The results now filter by vendor and sort by cost, consistency or price per hour, and models from this update are marked. Cohere Command R+ was retired by AWS before it could run.
Grok 4.6 joins in second place
xAI shipped Grok 4.6 after the main run, so it was tested on the same suite: 28 of 30 tasks, 12 of 14 hard tasks three times in a row, for an $8.91 bill.
Full field: 36 models, 30 tasks
The suite grew to 30 tasks, with 14 hard ones repeated three times for the leading models. About 1,400 graded runs, published with per-model cost and safety records.
First public results
Eleven models on the first version of the suite, roughly 400 runs, to check that the scores order the way practitioners expect before widening the field.
One shared practice workspace: 27 contacts, 9 companies, 4 lists and a funded credit balance, rebuilt fresh for every run.
A scripted teammate answers every confirmation prompt the same way, so runs are repeatable.
Results are checked in the database, not in the conversation. A convincing reply with wrong data scores zero.
Costs are what each model billed at provider list rates, with prompt caching applied where the provider supports it.
Claude 5 models ran without prompt caching, so their costs read about five times higher than a tuned setup.
Gemini 3.7 and 3.8 Flash are billed at their standard list rate; Google is discounting both by half until the end of 2026.
One GLM-5 run could not be scored, so GLM-5 counts 29 tasks and one hard task reads "of 36 models".
A few models run as the closest version we can access. Every number can be rechecked from the saved run logs.