Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,194 ms p50, 3,209 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Equivalent | Met | 97% | 1 | 1,268 ms 58,972 in / 1,742 out | 2,070 ms 58,972 in / 1,742 out | 540 ms 56,613 in / 1,463 out | -1,530 ms (-74%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 22 | Equivalent | Met | 97% | 1 | 2,378 ms 134,928 in / 5,111 out | 2,117 ms 134,928 in / 5,111 out | 571 ms 129,531 in / 4,293 out | -1,546 ms (-73%) | Evaluated | View details |
| 23 | Equivalent | Met | 97% | 2 | 1,040 ms 53,075 in / 1,698 out | 1,654 ms 53,075 in / 1,698 out | 602 ms 50,952 in / 1,426 out | -1,052 ms (-64%) | Evaluated | View details |
| 24 | Equivalent | Met | 97% | 1 | 305 ms 51,896 in / 1,966 out | 1,701 ms 51,896 in / 1,966 out | 633 ms 49,820 in / 1,651 out | -1,068 ms (-63%) | Evaluated | View details |
| 25 | Equivalent | Met | 97% | 1 | 3,885 ms 288,964 in / 8,935 out | 1,748 ms 288,964 in / 8,935 out | 664 ms 277,405 in / 7,505 out | -1,084 ms (-62%) | Evaluated | Hide details |
Case 25 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 26 | Equivalent | Met | 97% | 1 | 1,922 ms 153,328 in / 4,995 out | 1,795 ms 153,328 in / 4,995 out | 695 ms 147,195 in / 4,196 out | -1,100 ms (-61%) | Evaluated | View details |
| 27 | Equivalent | Met | 97% | 1 | 451 ms 150,261 in / 4,646 out | 1,842 ms 150,261 in / 4,646 out | 466 ms 144,251 in / 3,903 out | -1,376 ms (-75%) | Evaluated | View details |
| 28 | Equivalent | Met | 97% | 1 | 1,546 ms 81,146 in / 2,573 out | 1,889 ms 81,146 in / 2,573 out | 497 ms 77,900 in / 2,161 out | -1,392 ms (-74%) | Evaluated | View details |
| 29 | Equivalent | Met | 97% | 1 | 3,087 ms 259,478 in / 8,712 out | 1,936 ms 259,478 in / 8,712 out | 528 ms 249,099 in / 7,318 out | -1,408 ms (-73%) | Evaluated | View details |
| 30 | Not equivalent | Missed | 91% | 1 | 5,763 ms 253,581 in / 8,042 out | 1,983 ms 253,581 in / 8,042 out | 559 ms 243,438 in / 6,755 out | -1,424 ms (-72%) | Evaluated | View details |
Northstar support case 25: resolve the chat-resolver request with the account policy and cited help-center context.
gpt-5.5
Resolved chat-resolver case 25 with the correct policy, next step, and customer-safe explanation.
gpt-4o-mini
Cheaper candidate preserved the policy, action, and explanation for chat-resolver case 25.