Every model below ran the same 42 golden email chains through the same production prompt and was graded by the same scorer, so the only variable is the model. Accuracy, cost and latency are recorded in one pass, because a model that is cheaper and worse is not cheaper.
Incomplete comparison: claude-haiku-4.5 could not be reached on this run, so the head-to-head is OpenAI-only. Those rows are blank rather than zero.
| model | mean | passed every run | flaky | latency/call | cost/run |
|---|---|---|---|---|---|
| gpt-4o-mini | 100.0% | 100% | 0 | 0.97s | $0.0048 |
| gpt-4o | 100.0% | 100% | 0 | 0.86s | $0.0832 |
| claude-haiku-4.5 | unavailable this run | ||||