The Enginy model benchmark

The best AI models for sales, tested on real work.

44 models ran the same 30 sales jobs inside Enginy, 1,781 graded runs in all. If the work isn’t done right in the product, it doesn’t count.
Results

Which models get the work done

The top 14 of 44 models by share of tasks completed correctly. Filter by vendor, and hover a model for its numbers.

100%
Gemini 3.6 Flash
97%
Gemini 3.7 Flash
97%
Gemini 3.8 Flash
97%
Claude Fable 5.1
New
93%
Grok 4.6
93%
Claude Opus 5
93%
Gemini 3.5 Flash
90%
Grok 4.5
90%
GPT-6 Astra
New
90%
Claude Opus 4.8
87%
Claude Sonnet 4.6
87%
Claude Sonnet 5
83%
GPT-5.6 Terra
83%
GLM-5
added this update
Defaultwhat Enginy’s AI Chat runs on
Top 14 of 44 models. Hover a bar for details.
Value for money

The best results no longer cost the most

What a full test run costs, against how much work gets done. The models on the dashed line give you the most for your money.

0%20%40%60%80%100%$0.10$1$10$100Cost of a full test run, log scale. Up and to the left is better.Gemini 3.6 FlashClaude Fable 5.1†Claude Opus 5†Grok 4.5GPT-6 AstraClaude Sonnet 4.6GPT-5.6 LunaGPT-5.5Gemini 3.5 Flash-Litegpt-oss-120BGPT-5.4 Nanogpt-oss-20BGemma 3 27B
Green dots sit on the value line: nothing cheaper scores higher. Ringed dots were added this update.
† Claude 5 and Fable 5.1 costs are measured without prompt caching — a tuned setup bills roughly five times less.
About the benchmark

How the benchmark works

Enginy Sales Bench measures how well an AI model operates Enginy. Every model gets the same 30 real sales jobs in the same workspace: building lists, cleaning data, drafting campaigns. The hard ones run three times for the leading models. Work only counts when it’s completed correctly in the product, and every score comes with the model’s real bill.

Scores include repeat-run consistency on hard tasks and a pass/fail safety floor.

Costs are each model’s real bill for the full test. A person still reviews the work.

A few models run as the closest version we can access; every result can be rechecked from saved logs.

September 2026