AI test report — OpenAI and Claude

Complete and reviewed · 2026-10-05T12:05:56.334099+00:00

Index

Scope and models

279 scenario occurrences per model: 95 original UAT, 24 additions, 147 further acceptance/regression cases, and 13 latest E2E scenarios. 558 total model scenarios. Shared code tests are separate: 458 passed / 5 failed.

Component OpenAI variant Claude variant
Conversation + agenda openai/gpt-5.4-mini anthropic/claude-sonnet-4.6
Profile extraction openai/gpt-5.6-luna same
Narrow Jev selector active active
Legacy Jev routing/ticket/workflow interpreters off off
Simulator anthropic/claude-haiku-4.5 same
Judge anthropic/claude-sonnet-4.6 same

This compares the response/agenda model in the existing application, not an all-OpenAI stack against an all-Anthropic stack. Asynchronous conversation review remains on. No production or staging settings were changed.

Original UAT and later test blocks

Counts below are pass / partial / fail / blocked / error, preserving raw automatic judgments.

Block Expected per model OpenAI raw counts Claude raw counts
original-uat 95 44 / 29 / 19 / 3 / 0 (95 completed) 48 / 25 / 17 / 5 / 0 (95 completed)
uat-additions 24 7 / 6 / 10 / 1 / 0 (24 completed) 7 / 5 / 12 / 0 / 0 (24 completed)
knowledge 26 8 / 6 / 12 / 0 / 0 (26 completed) 14 / 3 / 9 / 0 / 0 (26 completed)
regression 38 18 / 7 / 13 / 0 / 0 (38 completed) 27 / 2 / 9 / 0 / 0 (38 completed)
feedback 14 7 / 3 / 4 / 0 / 0 (14 completed) 8 / 1 / 5 / 0 / 0 (14 completed)
ticket-precision 12 5 / 1 / 6 / 0 / 0 (12 completed) 4 / 2 / 6 / 0 / 0 (12 completed)
handoff-regression 4 2 / 0 / 1 / 1 / 0 (4 completed) 0 / 0 / 3 / 1 / 0 (4 completed)
workflow-router 12 11 / 0 / 1 / 0 / 0 (12 completed) 10 / 1 / 1 / 0 / 0 (12 completed)
flow-selector 23 19 / 1 / 2 / 1 / 0 (23 completed) 20 / 1 / 2 / 0 / 0 (23 completed)
response-policy 18 17 / 1 / 0 / 0 / 0 (18 completed) 18 / 0 / 0 / 0 / 0 (18 completed)
latest-e2e 13 13 / 0 / 0 / 0 / 0 (13 completed) 12 / 0 / 1 / 0 / 0 (13 completed)

OpenAI report · Claude report · Detailed evidence review

Model comparison

Claude has 17 more raw passes in this single run (168 versus 151), while OpenAI has lower measured first-token latency. Forty paired cases became passes under Claude, while 23 OpenAI passes did not remain passes under Claude. Stale expectations, adaptive follow-ups and judge/source gaps prevent treating this as a definitive model ranking.

Paired cases available: 279. Per-case CSV includes raw verdict transitions and median visitor first-token milliseconds per case. Adaptive follow-ups differ across models, so this is a paired case-specification comparison, not a replay of every identical message.

OpenAI verdict Claude verdict Cases
blocked blocked 3
blocked fail 1
blocked partial 1
blocked pass 1
fail fail 46
fail partial 8
fail pass 14
partial blocked 1
partial fail 6
partial partial 22
partial pass 25
pass blocked 2
pass fail 12
pass partial 9
pass pass 128

Raw totals

Variant Pass Partial Fail Blocked Error
openai 151 54 68 6 0
claude 168 40 65 6 0

Confirmed findings

The most important confirmed findings are shared consent-confirmation loops, a Claude duplicate attempt after a rejected sponsorship handoff, unredacted synthetic contact details in stored tool traces, and information questions that prematurely start qualification. See review notes for exact cases, evidence, stale expectations and source-backed corrections. No application fixes were mixed into the benchmark.

Priority Work Why
1 Correct telemetry masking for current SDK tool/generation attributes Synthetic contact details remain visible in stored traces.
1 Fix consent matching, known-email guard and explicit workflow switching Valid requests loop, ask for known details, or choose the wrong submission tool.
1 Prevent retry on a simple acknowledgement after failed handoff Claude made two rejected attempts in the scripted failure case.
2 Narrow past-event/privacy guards and contextual contact routing Employee headcount and email-delivery requests trigger unrelated refusals.
2 Answer information questions before qualification; preserve agenda constraints Policy questions, session types/access and unsolicited CTAs need attention.
2 Refresh judge ground truth from sources and current policies Historical prices/contact rules and unsupported invention allegations distort raw scores.

A response-model swap alone will not fix the shared deterministic defects. Preserve the comparative evidence and rerun the affected blocks after targeted fixes; do not treat these raw totals as a release gate or a causal model ranking.

Latency and streaming

Variant Measured turns Visitor first token p50 p90 Profile p50 Complete response p50 Multi-chunk turns Accepted local / attempts
openai 666 3.36s 5.57s 2.10s 3.49s 347 34 / 38
claude 683 3.84s 8.36s 2.15s 4.63s 361 34 / 39

Visitor first-token latency includes profile extraction. Local concurrent load and API buffering differ from a production browser. Multiple chunks demonstrate incremental delivery; deterministic short answers may use one chunk. HTTP 503 handoffs are rejected and never counted as delivered. E2E timings are kept in their original artifacts and excluded from acceptance percentiles.

Ordinary and operational checks

Full ordinary-test and operational checklist: 458 unit tests passed, five form-test fixtures failed under happy-dom; Studio build passed after setup recovery; web typecheck remains failing. Browser cookie gating and correct hotel/session-link behavior were checked. Inbox delivery, CRM processing and shared CMS mutation/ingestion checks were not performed.

Evidence and limitations

Targeted fixes and failure-only retest

Follow-up report: fixes are implemented and only the 37 selected prior application failures were rerun on their original failing model; the compiled follow-up records 16 pass / 10 partial / 11 fail, with source-backed review notes. Original raw verdicts above remain unchanged. The new FAQ freshness block is prepared separately and is not part of this rerun.