Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,194 ms p50, 3,209 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Equivalent | Met | 97% | 2 | 478 ms 49,537 in / 1,921 out | 1,640 ms 49,537 in / 1,921 out | 440 ms 47,556 in / 1,614 out | -1,200 ms (-73%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| Hide details |
Case 1 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured promptNorthstar support case 1: resolve the chat-resolver request with the account policy and cited help-center context. Original responsegpt-5.5 Resolved chat-resolver case 1 with the correct policy, next step, and customer-safe explanation. Candidate responsegpt-4o-mini Cheaper candidate preserved the policy, action, and explanation for chat-resolver case 1. |
| 2 | Equivalent | Met | 97% | 1 | 797 ms 144,128 in / 4,530 out | 1,687 ms 144,128 in / 4,530 out | 471 ms 138,363 in / 3,805 out | -1,216 ms (-72%) | Evaluated | View details |
| 3 | Equivalent | Met | 97% | 1 | 1,119 ms 241,786 in / 9,829 out | 1,734 ms 241,786 in / 9,829 out | 502 ms 232,115 in / 8,256 out | -1,232 ms (-71%) | Evaluated | View details |
| 4 | Equivalent | Met | 97% | 1 | 361 ms 86,807 in / 2,859 out | 1,781 ms 86,807 in / 2,859 out | 533 ms 83,335 in / 2,402 out | -1,248 ms (-70%) | Evaluated | View details |
| 5 | Equivalent | Met | 97% | 1 | 108 ms 28,307 in / 965 out | 1,828 ms 28,307 in / 965 out | 564 ms 27,175 in / 811 out | -1,264 ms (-69%) | Evaluated | View details |
| 6 | Equivalent | Met | 97% | 1 | 882 ms 57,793 in / 1,653 out | 1,875 ms 57,793 in / 1,653 out | 595 ms 55,481 in / 1,389 out | -1,280 ms (-68%) | Evaluated | View details |
| 7 | Equivalent | Met | 97% | 1 | 1,646 ms 56,613 in / 1,921 out | 1,922 ms 56,613 in / 1,921 out | 626 ms 54,348 in / 1,614 out | -1,296 ms (-67%) | Evaluated | View details |
| 8 | Equivalent | Met | 97% | 1 | 873 ms 54,254 in / 1,832 out | 1,969 ms 54,254 in / 1,832 out | 657 ms 52,084 in / 1,539 out | -1,312 ms (-67%) | Evaluated | View details |
| 9 | Equivalent | Met | 97% | 1 | 1,605 ms 131,862 in / 4,879 out | 2,016 ms 131,862 in / 4,879 out | 688 ms 126,588 in / 4,098 out | -1,328 ms (-66%) | Evaluated | View details |
| 10 | Equivalent | Met | 97% | 1 | 3,134 ms 128,795 in / 4,530 out | 2,063 ms 128,795 in / 4,530 out | 459 ms 123,643 in / 3,805 out | -1,604 ms (-78%) | Evaluated | View details |