Verdict
Quality matched on 97% of cases and the candidate is cheaper. Latency is informational during calibration.
Cases evaluated
Evaluation complete
Quality matches
97%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o-mini received the same frozen prompts in alternating order.
Replayed source p50
1,859 ms
gpt-5.5
Replayed candidate p50
556 ms
gpt-4o-mini
Paired median delta
-1,272 ms (-70%)
95% interval: -1,384 ms to -1,200 ms
Controlled p90
2,064 ms / 665 ms
Source / candidate
Production source context
1,194 ms p50, 3,209 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Equivalent | Met | 97% | 1 | 1,588 ms 122,662 in / 4,298 out | 2,110 ms 122,662 in / 4,298 out | 490 ms 117,756 in / 3,610 out | -1,620 ms (-77%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 12 | Equivalent | Met | 97% | 2 | 2,532 ms 283,067 in / 8,488 out | 1,647 ms 283,067 in / 8,488 out | 521 ms 271,744 in / 7,130 out | -1,126 ms (-68%) | Evaluated | View details |
| 13 | Equivalent | Met | 97% | 1 | 5,208 ms 277,169 in / 9,829 out | 1,694 ms 277,169 in / 9,829 out | 552 ms 266,082 in / 8,256 out | -1,142 ms (-67%) | Evaluated | View details |
| 14 | Equivalent | Met | 97% | 1 | 2,502 ms 265,375 in / 9,382 out | 1,741 ms 265,375 in / 9,382 out | 583 ms 254,760 in / 7,881 out | -1,158 ms (-67%) | Evaluated | View details |
| 15 | Equivalent | Met | 97% | 1 | 967 ms 79,259 in / 3,074 out | 1,788 ms 79,259 in / 3,074 out | 614 ms 76,089 in / 2,582 out | -1,174 ms (-66%) | Evaluated | View details |
| 16 | Equivalent | Met | 97% | 1 | 2,113 ms 77,372 in / 2,859 out | 1,835 ms 77,372 in / 2,859 out | 645 ms 74,277 in / 2,402 out | -1,190 ms (-65%) | Evaluated | Hide details |
Case 16 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 17 | Equivalent | Met | 97% | 1 | 954 ms 94,356 in / 2,716 out | 1,882 ms 94,356 in / 2,716 out | 676 ms 90,582 in / 2,281 out | -1,206 ms (-64%) | Evaluated | View details |
| 18 | Equivalent | Met | 97% | 1 | 377 ms 33,260 in / 1,045 out | 1,929 ms 33,260 in / 1,045 out | 447 ms 31,930 in / 878 out | -1,482 ms (-77%) | Evaluated | View details |
| 19 | Equivalent | Met | 97% | 1 | 887 ms 32,553 in / 965 out | 1,976 ms 32,553 in / 965 out | 478 ms 31,251 in / 811 out | -1,498 ms (-76%) | Evaluated | View details |
| 20 | Equivalent | Met | 97% | 1 | 371 ms 31,137 in / 1,153 out | 2,023 ms 31,137 in / 1,153 out | 509 ms 29,892 in / 969 out | -1,514 ms (-75%) | Evaluated | View details |
Northstar support case 16: resolve the chat-resolver request with the account policy and cited help-center context.
gpt-5.5
Resolved chat-resolver case 16 with the correct policy, next step, and customer-safe explanation.
gpt-4o-mini
Cheaper candidate preserved the policy, action, and explanation for chat-resolver case 16.