Evidence review

Both 279-case runs are complete. This evidence review supplements the raw verdicts without replacing them. It reviews critical failure categories, targeted workflow/policy cases and sampled tool traces; it is not a claim that every factual sentence was independently audited or that the raw totals are a release-readiness score.

Index

Verified working paths

Confirmed application issues

  1. Consent wording is too restrictive, shared by both models. Original UAT C01/C03 repeat the ticket confirmation instead of submitting after “Yes, please send it through.” OpenAI C17 and Claude C14 show related agenda confirmation loops. The pure matcher reproduces rejection of “Yes, please send it through” and “Yes, that’s all correct. Please send it”; “yes pls” and “Yes, please go ahead and send it” are accepted. See matcher evidence and implementation. A repeated confirmation is a failed handoff, not automatically a false claim of completed delivery.

  2. Claude retries a rejected sponsorship handoff after “thanks”. Latest E2E sponsorship-failure records two HTTP 503 attempts where exactly one was expected. The first explicit confirmation is valid; the later acknowledgement must not cause a retry. OpenAI passed this scenario. See Claude E2E results. Both attempts were local; neither was accepted or delivered.

  3. Stored trace PII is not fully redacted. A sampled Claude submitLead tool observation stores the synthetic name Jordan UAT and synthetic email directly. The associated generation stores tool arguments too. The code has an intended masking callback, but observed stored evidence does not meet O06. See trace evidence. Do not infer that every trace is unredacted, or that real visitor traces were inspected.

  4. Information requests can enter qualification prematurely. OpenAI original C78 asks whether tickets can be bought at the venue; C81 asks about upgrading an existing pass. Both get “What company are you with, and what is your role?” instead of answering the policy question. Original C20 on both models loops on pass selection when Miami options are unavailable. This is workflow behavior, not a content-only fix.

  5. Some answers drift away from the requested content. OpenAI original C24 returns an agenda session rather than articles/webinars; C26 returns editorial articles for reports; C85 substitutes agenda themes for attendee industries. These are distinct from fabricated source data.

  6. Invalid placeholder link in a response. OpenAI original C95 offers /events/unleash-paris/... as the sponsorship URL. It should use a real published destination. This was not called out by the raw partial verdict.

  7. Privacy guard can intercept legitimate delivery requests. OpenAI original C66: “Yes, absolutely - please send it to my email” receives “I can’t share your personal information here.” That is a false-positive privacy refusal during a consent flow.

  8. Exhibitor logistics contact routing is still incomplete. OpenAI regression R07 explicitly asks about submitting stand plans. The response recognizes exhibitor logistics but links support@unleash.ai. The current sponsor/exhibitor policy requires customersuccess@unleash.ai here. Unlike generic complaints/support cases, this is not an obsolete-email expectation.

  9. Speaker/host distinctions need attention. OpenAI regression R14 lists Stage 1 speakers in response to who hosts the stage. UAT C104 includes Brian Glaser's fireside under keynotes; the browser session flyout identifies that session as Fireside. Other roster differences still require current-source review.

  10. Employee count is mistaken for a past-event year. Feedback F04/F13 contain “2000 employees” in a current Paris ticket request. The pre-model policy sees 2000 as an old year and returns the past-event refusal, breaking the handoff. The pure function reproduces it; adding explicit 2026 prevents it. Policy matcher evidence also reproduces the “send it to my email” privacy false positive and stand-plan contact mismatch. These are application guard defects, independent of the response model.

  11. Known email is ignored by an agenda correction guard. W06 on both models starts with a complete saved profile and consent. “Please email the selected agenda to me” triggers the malformed-email question. Langfuse proves profile extraction retained the email; the agendaEmailPending branch matches please ... to without checking the saved address. See profile evidence and route. This is not the extractor forgetting the profile.

  12. Cross-flow tool choice can be wrong even after confirmation. OpenAI handoff H01 changes from an agenda discussion to ticket registration, but its final confirmation invokes sendAgendaByEmail with an empty session list instead of a delegate lead. The tool rejects it before the local receiver. OpenAI flow-selector W03 calls the sponsorship knowledge search and reports an unsuccessful handoff without calling submitLead. Claude H01 likewise remains in agenda-email mode after the summit/ticket enquiry, then re-asks for the email already supplied. These are routing/state/action failures, not evidence that the webhook server is down. Tool evidence.

  13. Agenda submission can fail before dispatch despite populated session links. Claude C107 calls sendAgendaByEmail with nine session URLs and receives ok:false; no capture is recorded. Claude handoff F14 also returns the generic submission error twice, with no receiver capture; the same scenario passed in its earlier feedback block. The exact failing guard is not exposed by the generic error. It requires contact/selection/consent validation diagnostics, not a claim of an external HTTP outage. optIn:false is the documented marketing default and is not itself evidence of missing request consent.

  14. A switched workflow can omit the complete pre-send summary. OpenAI flow-selector W10 switches SPEX to tickets, then sends on the next affirmative without showing the full saved contact summary. Correct lead type and one accepted local capture do not make the confirmation UX complete. Claude workflow-router W07 captures the corrected email correctly but omits it from the refreshed summary, preventing the visitor from checking that correction.

  15. Brand normalization produces awkward copy in an adversarial spelling case. OpenAI policy P07 says the legacy name is “UNLEASH WORLD,” not “UNLEASH WORLD.” Capitalization is enforced, but the rendered sentence contradicts itself.

  16. Free-form agenda responses can retain overlapping choices. Claude feedback F03 explicitly lists 10:00–10:30 and 10:25–10:50 in the same proposed plan, acknowledges the overlap, and suggests leaving early. The later plan is corrected and locally accepted for email, but the first response fails the non-overlapping one-day-plan requirement. Keep this distinct from clearly labeled optional alternatives outside a validated plan.

Grading and expectation issues

Ordinary tests and setup

Operational boundaries

All AI handoffs target loopback receivers. HTTP 2xx means accepted by the local test receiver, not delivered email or a Salesforce record. HTTP 503 is rejected, never successful delivery. Inbox timing, CRM processing, destructive ingestion, production kill-switch changes, and live CMS freshness mutations are not covered by this run. Published CMS content remains live/read-only; no shared configuration was changed. Legacy experimental Jev routing/ticket classifier ablations are not rerun: the active application endpoints cover the current selector and workflow suites instead.