The top 14 of 44 models by share of tasks completed correctly. Filter by vendor, and hover a model for its numbers.
What a full test run costs, against how much work gets done. The models on the dashed line give you the most for your money.
Enginy Sales Bench measures how well an AI model operates Enginy. Every model gets the same 30 real sales jobs in the same workspace: building lists, cleaning data, drafting campaigns. The hard ones run three times for the leading models. Work only counts when it’s completed correctly in the product, and every score comes with the model’s real bill.
Scores include repeat-run consistency on hard tasks and a pass/fail safety floor.
Costs are each model’s real bill for the full test. A person still reviews the work.
A few models run as the closest version we can access; every result can be rechecked from saved logs.