Verdict
Quality matched on 93% of cases. Review the mismatches before changing models.
Cases evaluated
Evaluation complete
Quality matches
93%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o received the same frozen prompts in alternating order.
Replayed source p50
2,079 ms
gpt-5.5
Replayed candidate p50
946 ms
gpt-4o
Paired median delta
-1,102 ms (-54%)
95% interval: -1,222 ms to -1,033 ms
Controlled p90
2,284 ms / 1,055 ms
Source / candidate
Production source context
936 ms p50, 2,640 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Equivalent | Met | 97% | 1 | 420 ms 47,571 in / 9,605 out | 2,290 ms 47,571 in / 9,605 out | 930 ms 45,668 in / 8,068 out | -1,360 ms (-59%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 22 | Equivalent | Met | 97% | 1 | 682 ms 137,739 in / 22,651 out | 2,337 ms 137,739 in / 22,651 out | 961 ms 132,229 in / 19,027 out | -1,376 ms (-59%) | Evaluated | View details |
| 23 | Equivalent | Met | 97% | 2 | 917 ms 232,449 in / 49,144 out | 1,874 ms 232,449 in / 49,144 out | 992 ms 223,151 in / 41,281 out | -882 ms (-47%) | Evaluated | View details |
| 24 | Equivalent | Met | 97% | 1 | 274 ms 83,033 in / 14,296 out | 1,921 ms 83,033 in / 14,296 out | 1,023 ms 79,712 in / 12,009 out | -898 ms (-47%) | Evaluated | Hide details |
Case 24 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 25 | Equivalent | Met | 97% | 1 | 1,069 ms 27,245 in / 4,825 out | 1,968 ms 27,245 in / 4,825 out | 1,054 ms 26,155 in / 4,053 out | -914 ms (-46%) | Evaluated | View details |
| 26 | Equivalent | Met | 97% | 1 | 305 ms 44,327 in / 9,382 out | 2,015 ms 44,327 in / 9,382 out | 1,085 ms 42,554 in / 7,881 out | -930 ms (-46%) | Evaluated | View details |
| 27 | Equivalent | Met | 97% | 1 | 451 ms 129,306 in / 22,070 out | 2,062 ms 129,306 in / 22,070 out | 856 ms 124,134 in / 18,539 out | -1,206 ms (-58%) | Evaluated | View details |
| 28 | Equivalent | Met | 97% | 1 | 5,763 ms 216,231 in / 48,027 out | 2,109 ms 216,231 in / 48,027 out | 887 ms 207,582 in / 40,343 out | -1,222 ms (-58%) | Evaluated | View details |
| 29 | Not equivalent | Missed | 91% | 1 | 2,351 ms 77,843 in / 13,939 out | 2,156 ms 77,843 in / 13,939 out | 918 ms 74,729 in / 11,709 out | -1,238 ms (-57%) | Evaluated | View details |
| 30 | Not equivalent | Missed | 91% | 1 | 992 ms 32,435 in / 5,897 out | 2,203 ms 32,435 in / 5,897 out | 949 ms 31,138 in / 4,953 out | -1,254 ms (-57%) | Evaluated | View details |
Northstar support case 24: resolve the conversation-recap request with the account policy and cited help-center context.
gpt-5.5
Resolved conversation-recap case 24 with the correct policy, next step, and customer-safe explanation.
gpt-4o
Cheaper candidate preserved the policy, action, and explanation for conversation-recap case 24.