Verdict
Quality matched on 93% of cases. Review the mismatches before changing models.
Cases evaluated
Evaluation complete
Quality matches
93%
Est. monthly savings
Projected from matching events observed over 30 days
Status
The experiment finished and its verdict is available.
Controlled replay compares both models from the same worker. Production latency is shown separately as context and is not subtracted from candidate replay latency.
gpt-5.5 and gpt-4o received the same frozen prompts in alternating order.
Replayed source p50
2,079 ms
gpt-5.5
Replayed candidate p50
946 ms
gpt-4o
Paired median delta
-1,102 ms (-54%)
95% interval: -1,222 ms to -1,033 ms
Controlled p90
2,284 ms / 1,055 ms
Source / candidate
Production source context
936 ms p50, 2,640 ms p90
30 captured measurements from the user's environment
Controlled coverage
30 successful pairs
Only complete source/candidate pairs enter the comparison (provider-completion-v1)
Replay failures
The evaluation compares structured content canonically and reports presentation adherence as separate evidence unless strict formatting is enabled for the feature.
This run's encrypted evidence expires at the date shown above, 14 days after the run was created. The verdict summary and savings estimate remain in experiment history after that run-specific evidence is deleted.
Review the case results before opening sensitive prompt and response evidence. Opening a case is audited.
| Case | Result | Format | Evaluation confidence | Attempts | Production source | Replayed source | Candidate | Replay delta | Status | Case details |
|---|---|---|---|---|---|---|---|---|---|---|
| 11 | Equivalent | Met | 97% | 1 | 365 ms 118,062 in / 22,070 out | 2,330 ms 118,062 in / 22,070 out | 880 ms 113,340 in / 18,539 out | -1,450 ms (-62%) | Evaluated |
0 source, 0 candidate
Failures are excluded from percentiles and retained as reliability evidence
| View details |
| 12 | Equivalent | Met | 97% | 2 | 371 ms 28,543 in / 5,763 out | 1,867 ms 28,543 in / 5,763 out | 911 ms 27,401 in / 4,841 out | -956 ms (-51%) | Evaluated | Hide details |
Case 12 evidenceFull captured evidence for this scored comparison. EquivalentFormat metscored97% evaluation confidenceJudge: gemini-3.1-pro-preview Evaluation notesThe candidate preserved the original answer's policy, resolution, and customer-safe guidance. Format evidenceThe response preserved the expected support-answer structure. Captured prompt | ||||||||||
| 13 | Equivalent | Met | 97% | 1 | 5,611 ms 254,072 in / 48,027 out | 1,914 ms 254,072 in / 48,027 out | 942 ms 243,909 in / 40,343 out | -972 ms (-51%) | Evaluated | View details |
| 14 | Equivalent | Met | 97% | 1 | 1,268 ms 54,058 in / 8,712 out | 1,961 ms 54,058 in / 8,712 out | 973 ms 51,896 in / 7,318 out | -988 ms (-50%) | Evaluated | View details |
| 15 | Equivalent | Met | 97% | 1 | 2,286 ms 70,924 in / 13,939 out | 2,008 ms 70,924 in / 13,939 out | 1,004 ms 68,087 in / 11,709 out | -1,004 ms (-50%) | Evaluated | View details |
| 16 | Equivalent | Met | 97% | 1 | 2,378 ms 123,684 in / 25,555 out | 2,055 ms 123,684 in / 25,555 out | 1,035 ms 118,737 in / 21,466 out | -1,020 ms (-50%) | Evaluated | View details |
| 17 | Equivalent | Met | 97% | 1 | 963 ms 29,840 in / 5,897 out | 2,102 ms 29,840 in / 5,897 out | 1,066 ms 28,646 in / 4,953 out | -1,036 ms (-49%) | Evaluated | View details |
| 18 | Equivalent | Met | 97% | 1 | 3,885 ms 264,884 in / 44,676 out | 2,149 ms 264,884 in / 44,676 out | 837 ms 254,289 in / 37,528 out | -1,312 ms (-61%) | Evaluated | View details |
| 19 | Equivalent | Met | 97% | 1 | 1,546 ms 74,384 in / 12,867 out | 2,196 ms 74,384 in / 12,867 out | 868 ms 71,409 in / 10,808 out | -1,328 ms (-60%) | Evaluated | View details |
| 20 | Equivalent | Met | 97% | 1 | 635 ms 31,137 in / 5,495 out | 2,243 ms 31,137 in / 5,495 out | 899 ms 29,892 in / 4,616 out | -1,344 ms (-60%) | Evaluated | View details |
Northstar support case 12: resolve the conversation-recap request with the account policy and cited help-center context.
gpt-5.5
Resolved conversation-recap case 12 with the correct policy, next step, and customer-safe explanation.
gpt-4o
Cheaper candidate preserved the policy, action, and explanation for conversation-recap case 12.