TM GO · model selection

Which model should be making the call?

Every model below ran the same 42 golden email chains through the same production prompt and was graded by the same scorer, so the only variable is the model. Accuracy, cost and latency are recorded in one pass, because a model that is cheaper and worse is not cheaper.

Incomplete comparison: claude-haiku-4.5 could not be reached on this run, so the head-to-head is OpenAI-only. Those rows are blank rather than zero.

gpt-4o costs 17.3× what gpt-4o-mini costs and scores +0.0 points for it.

Accuracy

mean score, 42 cases
gpt-4o-mini100.0%
gpt-4o100.0%
claude-haiku-4.5unavailable
Bars scaled to the best score.

Cost

USD per full golden run
gpt-4o-mini$0.0048
gpt-4o$0.0832
claude-haiku-4.5unavailable
Bars scaled to the most expensive.

Latency

average seconds per call
gpt-4o-mini0.97s
gpt-4o0.86s
claude-haiku-4.5unavailable
Bars scaled to the slowest.
modelmeanpassed every run flakylatency/callcost/run
gpt-4o-mini100.0%100%00.97s$0.0048
gpt-4o100.0%100%00.86s$0.0832
claude-haiku-4.5unavailable this run