AI test report — OpenAI and Claude
Complete and reviewed · 2026-10-05T12:05:56.334099+00:00
Index
Scope and models
279 scenario occurrences per model: 95 original UAT, 24 additions, 147 further acceptance/regression cases, and 13 latest E2E scenarios. 558 total model scenarios. Shared code tests are separate: 458 passed / 5 failed.
| Component |
OpenAI variant |
Claude variant |
| Conversation + agenda |
openai/gpt-5.4-mini |
anthropic/claude-sonnet-4.6 |
| Profile extraction |
openai/gpt-5.6-luna |
same |
| Narrow Jev selector |
active |
active |
| Legacy Jev routing/ticket/workflow interpreters |
off |
off |
| Simulator |
anthropic/claude-haiku-4.5 |
same |
| Judge |
anthropic/claude-sonnet-4.6 |
same |
This compares the response/agenda model in the existing application, not an all-OpenAI stack against an all-Anthropic stack. Asynchronous conversation review remains on. No production or staging settings were changed.
Original UAT and later test blocks
Counts below are pass / partial / fail / blocked / error, preserving raw automatic judgments.
| Block |
Expected per model |
OpenAI raw counts |
Claude raw counts |
| original-uat |
95 |
44 / 29 / 19 / 3 / 0 (95 completed) |
48 / 25 / 17 / 5 / 0 (95 completed) |
| uat-additions |
24 |
7 / 6 / 10 / 1 / 0 (24 completed) |
7 / 5 / 12 / 0 / 0 (24 completed) |
| knowledge |
26 |
8 / 6 / 12 / 0 / 0 (26 completed) |
14 / 3 / 9 / 0 / 0 (26 completed) |
| regression |
38 |
18 / 7 / 13 / 0 / 0 (38 completed) |
27 / 2 / 9 / 0 / 0 (38 completed) |
| feedback |
14 |
7 / 3 / 4 / 0 / 0 (14 completed) |
8 / 1 / 5 / 0 / 0 (14 completed) |
| ticket-precision |
12 |
5 / 1 / 6 / 0 / 0 (12 completed) |
4 / 2 / 6 / 0 / 0 (12 completed) |
| handoff-regression |
4 |
2 / 0 / 1 / 1 / 0 (4 completed) |
0 / 0 / 3 / 1 / 0 (4 completed) |
| workflow-router |
12 |
11 / 0 / 1 / 0 / 0 (12 completed) |
10 / 1 / 1 / 0 / 0 (12 completed) |
| flow-selector |
23 |
19 / 1 / 2 / 1 / 0 (23 completed) |
20 / 1 / 2 / 0 / 0 (23 completed) |
| response-policy |
18 |
17 / 1 / 0 / 0 / 0 (18 completed) |
18 / 0 / 0 / 0 / 0 (18 completed) |
| latest-e2e |
13 |
13 / 0 / 0 / 0 / 0 (13 completed) |
12 / 0 / 1 / 0 / 0 (13 completed) |
OpenAI report · Claude report · Detailed evidence review
Model comparison
Claude has 17 more raw passes in this single run (168 versus 151), while OpenAI has lower measured first-token latency. Forty paired cases became passes under Claude, while 23 OpenAI passes did not remain passes under Claude. Stale expectations, adaptive follow-ups and judge/source gaps prevent treating this as a definitive model ranking.
Paired cases available: 279. Per-case CSV includes raw verdict transitions and median visitor first-token milliseconds per case. Adaptive follow-ups differ across models, so this is a paired case-specification comparison, not a replay of every identical message.
| OpenAI verdict |
Claude verdict |
Cases |
| blocked |
blocked |
3 |
| blocked |
fail |
1 |
| blocked |
partial |
1 |
| blocked |
pass |
1 |
| fail |
fail |
46 |
| fail |
partial |
8 |
| fail |
pass |
14 |
| partial |
blocked |
1 |
| partial |
fail |
6 |
| partial |
partial |
22 |
| partial |
pass |
25 |
| pass |
blocked |
2 |
| pass |
fail |
12 |
| pass |
partial |
9 |
| pass |
pass |
128 |
Raw totals
| Variant |
Pass |
Partial |
Fail |
Blocked |
Error |
| openai |
151 |
54 |
68 |
6 |
0 |
| claude |
168 |
40 |
65 |
6 |
0 |
Confirmed findings
The most important confirmed findings are shared consent-confirmation loops, a Claude duplicate attempt after a rejected sponsorship handoff, unredacted synthetic contact details in stored tool traces, and information questions that prematurely start qualification. See review notes for exact cases, evidence, stale expectations and source-backed corrections. No application fixes were mixed into the benchmark.
Recommended next fixes
| Priority |
Work |
Why |
| 1 |
Correct telemetry masking for current SDK tool/generation attributes |
Synthetic contact details remain visible in stored traces. |
| 1 |
Fix consent matching, known-email guard and explicit workflow switching |
Valid requests loop, ask for known details, or choose the wrong submission tool. |
| 1 |
Prevent retry on a simple acknowledgement after failed handoff |
Claude made two rejected attempts in the scripted failure case. |
| 2 |
Narrow past-event/privacy guards and contextual contact routing |
Employee headcount and email-delivery requests trigger unrelated refusals. |
| 2 |
Answer information questions before qualification; preserve agenda constraints |
Policy questions, session types/access and unsolicited CTAs need attention. |
| 2 |
Refresh judge ground truth from sources and current policies |
Historical prices/contact rules and unsupported invention allegations distort raw scores. |
A response-model swap alone will not fix the shared deterministic defects. Preserve the comparative evidence and rerun the affected blocks after targeted fixes; do not treat these raw totals as a release gate or a causal model ranking.
Latency and streaming
| Variant |
Measured turns |
Visitor first token p50 |
p90 |
Profile p50 |
Complete response p50 |
Multi-chunk turns |
Accepted local / attempts |
| openai |
666 |
3.36s |
5.57s |
2.10s |
3.49s |
347 |
34 / 38 |
| claude |
683 |
3.84s |
8.36s |
2.15s |
4.63s |
361 |
34 / 39 |
Visitor first-token latency includes profile extraction. Local concurrent load and API buffering differ from a production browser. Multiple chunks demonstrate incremental delivery; deterministic short answers may use one chunk. HTTP 503 handoffs are rejected and never counted as delivered. E2E timings are kept in their original artifacts and excluded from acceptance percentiles.
Ordinary and operational checks
Full ordinary-test and operational checklist: 458 unit tests passed, five form-test fixtures failed under happy-dom; Studio build passed after setup recovery; web typecheck remains failing. Browser cookie gating and correct hotel/session-link behavior were checked. Inbox delivery, CRM processing and shared CMS mutation/ingestion checks were not performed.
Evidence and limitations
Targeted fixes and failure-only retest
Follow-up report: fixes are implemented and only the 37 selected prior application failures were rerun on their original failing model; the compiled follow-up records 16 pass / 10 partial / 11 fail, with source-backed review notes. Original raw verdicts above remain unchanged. The new FAQ freshness block is prepared separately and is not part of this rerun.