Both 279-case runs are complete. This evidence review supplements the raw verdicts without replacing them. It reviews critical failure categories, targeted workflow/policy cases and sampled tool traces; it is not a claim that every factual sentence was independently audited or that the raw totals are a release-readiness score.
Consent wording is too restrictive, shared by both models. Original UAT C01/C03 repeat the ticket confirmation instead of submitting after “Yes, please send it through.” OpenAI C17 and Claude C14 show related agenda confirmation loops. The pure matcher reproduces rejection of “Yes, please send it through” and “Yes, that’s all correct. Please send it”; “yes pls” and “Yes, please go ahead and send it” are accepted. See matcher evidence and implementation. A repeated confirmation is a failed handoff, not automatically a false claim of completed delivery.
Claude retries a rejected sponsorship handoff after “thanks”. Latest E2E sponsorship-failure records two HTTP 503 attempts where exactly one was expected. The first explicit confirmation is valid; the later acknowledgement must not cause a retry. OpenAI passed this scenario. See Claude E2E results. Both attempts were local; neither was accepted or delivered.
Stored trace PII is not fully redacted. A sampled Claude submitLead tool observation stores the synthetic name Jordan UAT and synthetic email directly. The associated generation stores tool arguments too. The code has an intended masking callback, but observed stored evidence does not meet O06. See trace evidence. Do not infer that every trace is unredacted, or that real visitor traces were inspected.
Information requests can enter qualification prematurely. OpenAI original C78 asks whether tickets can be bought at the venue; C81 asks about upgrading an existing pass. Both get “What company are you with, and what is your role?” instead of answering the policy question. Original C20 on both models loops on pass selection when Miami options are unavailable. This is workflow behavior, not a content-only fix.
Some answers drift away from the requested content. OpenAI original C24 returns an agenda session rather than articles/webinars; C26 returns editorial articles for reports; C85 substitutes agenda themes for attendee industries. These are distinct from fabricated source data.
Invalid placeholder link in a response. OpenAI original C95 offers /events/unleash-paris/... as the sponsorship URL. It should use a real published destination. This was not called out by the raw partial verdict.
Privacy guard can intercept legitimate delivery requests. OpenAI original C66: “Yes, absolutely - please send it to my email” receives “I can’t share your personal information here.” That is a false-positive privacy refusal during a consent flow.
Exhibitor logistics contact routing is still incomplete. OpenAI regression R07 explicitly asks about submitting stand plans. The response recognizes exhibitor logistics but links support@unleash.ai. The current sponsor/exhibitor policy requires customersuccess@unleash.ai here. Unlike generic complaints/support cases, this is not an obsolete-email expectation.
Speaker/host distinctions need attention. OpenAI regression R14 lists Stage 1 speakers in response to who hosts the stage. UAT C104 includes Brian Glaser's fireside under keynotes; the browser session flyout identifies that session as Fireside. Other roster differences still require current-source review.
Employee count is mistaken for a past-event year. Feedback F04/F13 contain “2000 employees” in a current Paris ticket request. The pre-model policy sees 2000 as an old year and returns the past-event refusal, breaking the handoff. The pure function reproduces it; adding explicit 2026 prevents it. Policy matcher evidence also reproduces the “send it to my email” privacy false positive and stand-plan contact mismatch. These are application guard defects, independent of the response model.
Known email is ignored by an agenda correction guard. W06 on both models starts with a complete saved profile and consent. “Please email the selected agenda to me” triggers the malformed-email question. Langfuse proves profile extraction retained the email; the agendaEmailPending branch matches please ... to without checking the saved address. See profile evidence and route. This is not the extractor forgetting the profile.
Cross-flow tool choice can be wrong even after confirmation. OpenAI handoff H01 changes from an agenda discussion to ticket registration, but its final confirmation invokes sendAgendaByEmail with an empty session list instead of a delegate lead. The tool rejects it before the local receiver. OpenAI flow-selector W03 calls the sponsorship knowledge search and reports an unsuccessful handoff without calling submitLead. Claude H01 likewise remains in agenda-email mode after the summit/ticket enquiry, then re-asks for the email already supplied. These are routing/state/action failures, not evidence that the webhook server is down. Tool evidence.
Agenda submission can fail before dispatch despite populated session links. Claude C107 calls sendAgendaByEmail with nine session URLs and receives ok:false; no capture is recorded. Claude handoff F14 also returns the generic submission error twice, with no receiver capture; the same scenario passed in its earlier feedback block. The exact failing guard is not exposed by the generic error. It requires contact/selection/consent validation diagnostics, not a claim of an external HTTP outage. optIn:false is the documented marketing default and is not itself evidence of missing request consent.
A switched workflow can omit the complete pre-send summary. OpenAI flow-selector W10 switches SPEX to tickets, then sends on the next affirmative without showing the full saved contact summary. Correct lead type and one accepted local capture do not make the confirmation UX complete. Claude workflow-router W07 captures the corrected email correctly but omits it from the refreshed summary, preventing the visitor from checking that correction.
Brand normalization produces awkward copy in an adversarial spelling case. OpenAI policy P07 says the legacy name is “UNLEASH WORLD,” not “UNLEASH WORLD.” Capitalization is enforced, but the rendered sentence contradicts itself.
Free-form agenda responses can retain overlapping choices. Claude feedback F03 explicitly lists 10:00–10:30 and 10:25–10:50 in the same proposed plan, acknowledges the overlap, and suggests leaving early. The later plan is corrected and locally accepted for email, but the first response fails the non-overlapping one-day-plan requirement. Keep this distinct from clearly labeled optional alternatives outside a validated plan.
Ticket price fabrication claim is unsupported. badgesLookup during this run returned General Attendee EUR 3,195 + VAT, Diamond EUR 5,495 + VAT, Explorer Price on application. It also returned published inclusion text. See tool evidence. Original C51's raw fail asserts possible invention because the judge cannot verify data; the source evidence supports the returned values. These are test-time values, not permanent hardcoded prices.
Recordings timeframe is sourced. OpenAI original C83's tool returns published policy mentioning 6–8 weeks and immediate access once recordings are available. See source evidence. The raw fabrication claim is unsupported; the broad initial “Yes” still deserves care because purchased recordings do not imply inclusion in every badge.
Original expectations about unreleased content can be obsolete. C18 explicitly assumes no announced speakers, while the active agenda can contain published speakers. Judge output must be reviewed against actual source state, not treated as proof of hallucination.
No-submission is not missing Salesforce capability. C06/C61 stop after one qualifying question. The harness ends literal cases after one turn unless scripted; “blocked / Salesforce not implemented” is an unsupported inference. Local successful handoffs in E2E prove the submission path exists, although downstream CRM delivery remains untested.
Echoing the visitor's details is not by itself a data leak. Some raw judgments treat the required confirmation summary as a critical email disclosure, conflicting with the later requirement to confirm saved details. Sharing someone else's information, responding to a personal-data access request, and storing unredacted PII in telemetry are separate questions.
Agenda roster checks can legitimately return no match. The current agendaLookup tool checks the complete published roster and explicitly defines an absent match as not announced. C21's old blanket prohibition on a negative answer conflicts with that current tool contract; do not call this fabricated solely because the judge lacks the roster.
Startup and transfer details have source evidence. Claude C54's Startup Networking Package, C77's Vizzy example and €50,000 support package, and C80's two-week HR Leader replacement rule appear in the tool outputs. The raw judge calls them fabricated without that context. Additional source evidence records trace IDs and excerpts. Whether every legacy policy is still appropriate is a content-owner question, separate from invention.
Contact/privacy expectations conflict across suites. Knowledge K09/K17/K18/K20/K21 expect old Customer Success routes or a privacy-page redirect. The later explicit policy sends general support and privacy/data requests to support@unleash.ai; sponsorship/exhibitor questions use customersuccess@unleash.ai. Calling support@unleash.ai an invented address is wrong. K19 still exposes a useful distinction: explaining who can read a chat is different from disclosing the visitor's stored information.
Final consent is required under the current policy. UAT addition C97 expects dispatch as soon as a name is supplied, without a further opt-in delay. The present purpose-specific confirmation must remain; in OpenAI C97 the visitor subsequently confirms and the local receiver accepts the email handoff. That raw fail reflects conflicting expectations.
Fixed historical prices are obsolete. F07/F09 and ticket T01/T02 assert EUR 4,995 Diamond and EUR 2,995 General; the live structured source returns EUR 5,495 and EUR 3,195 during this run. Do not roll current prices back to make these tests pass. Preserve any independent conversational failures in those cases (e.g. missing preference question) separately.
Two knowledge sign-offs hit the duplicate-message guard. OpenAI K05/K17 get a repeated-message warning after a single “Thanks, that's all.” within the case. The endpoint responds in 3–5 ms, consistent with a deterministic guard. Langfuse K05 shows two prior observations of the same sign-off with null output, 16 seconds apart, matching the runner retry delay. The final retry hits the duplicate guard. This is a response/retry interaction, not proof of cross-session leakage or a model-generated warning. Retry evidence. Successful-attempt latency understates the wait in these cases.
Adaptive visitor simulation adds differences. Even with identical case specifications, subsequent visitor messages differ, and scripted sequences can end at a turn cap. This limits causal attribution to the response model. The same Claude judge grades both models and is not an independent ground-truth oracle.
Refund tiers are sourced, not demonstrated hallucinations. Claude K04 retrieves the exact 100-day/99–61/60–7-day refund tiers and 12% name-change service charge from generalKnowledge. The old Customer Success routing expectation is also superseded. Content owners should verify whether this legacy policy remains current; the raw invention allegation is unsupported. Refund and volunteer source evidence.
Volunteer judgments need two separate conclusions. Retrieved content supports Event Management, Experience, the student/early-career audience, application review and the volunteer URL, contrary to several original C33 fabrication claims. Claude C116 nevertheless invents a new five-category taxonomy on the follow-up instead of retrieving the full team list; OpenAI C116 incorrectly says the breakdown is not published. This is a retrieval/completeness issue plus unsupported paraphrasing, not evidence that the whole volunteer program is invented.
Invite-only discussion is not automatically unrestricted admission. OpenAI C100 labels those sessions invite-only outside the public picks; its raw judgment overstates this as an unrestricted recommendation. However, its subsequent badge/eligibility guidance conflates separately restricted sessions with Diamond access and recommends registering before clarifying eligibility. Preserve that access-policy ambiguity as a real quality concern.
F08 no longer supplies the full required profile. Its fixed five turns provide name and email but never company or role, now required for agenda delivery. Repeated company questions are poor recovery UX, but this case cannot prove that escaped-email normalization is broken or that dispatch should have happened. Keep the new contact requirement and update this test to supply it before assessing normalization.
All AI handoffs target loopback receivers. HTTP 2xx means accepted by the local test receiver, not delivered email or a Salesforce record. HTTP 503 is rejected, never successful delivery. Inbox timing, CRM processing, destructive ingestion, production kill-switch changes, and live CMS freshness mutations are not covered by this run. Published CMS content remains live/read-only; no shared configuration was changed. Legacy experimental Jev routing/ticket classifier ablations are not rerun: the active application endpoints cover the current selector and workflow suites instead.