Reviewed findings — targeted AI retest
Main report
Index
Verified application fixes
| Behavior |
Evidence |
Interpretation |
| Natural ticket-send confirmation |
C01 and C03, both models, first retest |
Exactly one HTTP 200 local lead capture follows the displayed contact/request summary and affirmative reply. C01 remains raw partial because of its qualification-order/free-HR-pass expectations; the consent-loop defect itself is fixed in these conversations. |
| Event clarification |
C03, both models |
Asks which event before proceeding, then collects details and confirms the Paris enquiry. |
| Saved agenda address and request summary |
W06 in workflow-router and flow-selector, round 2 |
Contact summary is shown first; the next confirmation produces one agenda-email capture, rather than a lead-type substitute. Selected plan is retained, and the alternative/conflicting sessions are excluded from the emailed plan. |
| Full agenda-email acceptance |
Claude C14, round 2 |
“Yes, please email it to me” keeps email qualification active while contact details are collected. The app supplies the summary; confirmation produces a single HTTP 200 agenda-email capture. Remaining raw partial is discussed below. |
| Browsing list converted into a valid personal plan |
Claude C107 and F03, first retest |
Both raw pass. The generated personal plan/email now meets the checked non-overlap constraints. This is not a promise that all possible interests/source data always produce a valid plan. |
| Ticket navigation |
C108, both models, round 2 |
The first ticket question receives registration/badge links; travel receives the internal travel page and direct hotel booking partner link. No contact-capture flow is started. |
| Existing-pass upgrade |
OpenAI C81, round 2; Claude C81, first retest |
Information is returned instead of a new-sales qualification question. Published policy gives case-by-case upgrades, subject to availability. The response still uses the source's department name “Customer Success” for ticket help; current app policy chooses Support, so department wording remains inconsistent. |
| Employee count versus historical edition |
F04, both models |
Headcount of 2,000 is no longer mistaken for an event year. Ticket request is summarized and confirmed. |
| Failed response retry |
OpenAI R26 |
Raw pass; empty failed HTTP attempts no longer count as previously delivered answers. |
| Stage host versus speaker list |
OpenAI R14 |
Raw pass; identifies David Green as host instead of substituting all stage speakers. |
| Exhibitor contact routing |
Claude R07 |
Raw pass; gives the exhibitor manual and customersuccess@unleash.ai. OpenAI still fails this behavior. |
| Failed sponsorship acknowledgement |
Claude sponsorship-failure |
One forced HTTP 503 attempt, followed by “thanks”; no second attempt, no false success claim, and Support fallback. The capture is rejected, not accepted delivery. |
| Persisted tool-argument email masking |
Synthetic post-fix Langfuse submitLead observation |
Stored tool arguments show [email redacted]. Names, company and role remain visible. Broad phone masking also masks timestamp/identifier fragments; full anonymisation is not established. Evidence. |
Remaining application defects
| Case |
Finding |
Evidence |
| H01, Claude |
Agenda-to-registration switch still enters agenda-email collection. A supplied email is subsequently requested again, and no delegate enquiry is captured. This remains a genuine workflow/state failure. |
Transcript |
| R07, OpenAI |
Stand-plan submission question receives the sponsorship opportunities page rather than the exhibitor manual/Customer Success contact. Answering as information avoids qualification but does not fix the wrong destination. |
Latest transcript |
| C104, both models |
Non-keynote formats are still included in keynote lists. Brian Glaser's session was identified as Fireside in the baseline browser/source review. Blanket reliance on the old “six speakers” roster is obsolete, but that does not excuse current format confusion. Prompt-only instructions were insufficient. Current session-format/marker reconciliation remains necessary. |
OpenAI, Claude, baseline source review |
| C116, both models |
Follow-up volunteer-role answers omit published teams. OpenAI also says there is no fixed published role list. The current retrieval/answer handling still fails to assemble the requested published taxonomy. |
OpenAI, Claude |
| R01/R02, OpenAI |
Free-ticket queries now receive information instead of qualification, but omit the expected HR-professionals/group/sponsor routes. This remains incomplete retrieval/answer coverage. Eligibility labels must be reconciled with current sources before promising complimentary access. |
R01, R02 |
| C78, OpenAI |
Qualification detour is removed, but the answer leads with an affirmative onsite-purchase claim. Collected tool observations contain badge pricing; no onsite FAQ lookup was present. Published FAQ uses conditional availability, so the answer should preserve that uncertainty. Pricing figures themselves are verified, not invented. |
Transcript, tool evidence |
There are also unfinished refinements among raw partial cases: C24 returns five articles without the expected webinar; C26 mixes reports/report-led articles with other content; C85 does not establish its industry list against approved audience data. Relative URLs are supported by the app and are not inherently invalid, but this retest did not exhaustively resolve every recommended content link. An unrelated personalised-agenda CTA still appears after some factual answers.
Content, retrieval and stale expectations
C78, Claude: the raw fail calls 3,195/5,495 pricing and conditional onsite registration invented. Stored Langfuse tool evidence contradicts that allegation: badgesLookup returned those exact structured displayPrice values, and generalKnowledge returned “On-site registration may be available … if capacity allows.” The answer follows that condition. Keep the raw fail, but classify this as a judge evidence gap rather than fabricated prices. This is not a blanket approval of every extra sentence in the answer. Source evidence.
C98, OpenAI: the original privacy false-positive is corrected in the retest. Its remaining raw fail concerns “personalised” and “programmes”; the current approved app-owned CTA uses “personalised.” This is a stale/conflicting wording criterion and must not prompt reverting the approved UI wording. The Claude occurrence was excluded initially because its baseline failure was solely wording.
W06, Claude: both blocks are raw partial because the judge wants the agenda reprinted immediately before confirmation. The seeded plan was already visible; the app shows the recipient/contact summary and captures exactly one confirmed agenda-email request with selected session URLs preserved. This is a summary-display expectation issue, not another send loop.
C14, Claude: round-two email handoff is accepted locally after a displayed summary and confirmation. The remaining raw partial penalizes echoing the email, even though confirming the recipient is now an explicit product requirement. It also notes no explicit lunch/networking/exhibition schedule; the app only provides a general footer for those. Do not call actual inbox arrival proven: the judge's “O02 satisfied” means a captured HTTP 200 request here.
C01 and Claude R01/R02: raw partials still apply an older qualification order/“HR professionals programme” naming requirement. The displayed confirmation and accepted handoff establish the consent fix, not eligibility correctness. Sources still contain a HR-professionals FAQ route alongside current Explorer/badge content; this requires maintained business/source ownership, not invented universal free-pass criteria.
C104: its August speaker-name ground truth is not a current complete roster certification. Evaluate the actual format mistake separately from the outdated exact roster. The source-content/keynote-marker discrepancy remains unresolved.
The automatic judge sees conversation/capture ground truth but not all retrieval tool outputs. Absence of sources in its injected context is not proof of fabrication. These notes preserve each raw verdict and add evidence; no raw labels were rewritten green.
External and operational limits
No new HTTP 402 billing blocker occurred in these runs. The initial acceptance setup failure was local JSON shape, not credit exhaustion, and happened before those acceptance calls ran. Its archived stale copies do not count as retest results.
The sponsorship failure case deliberately returns HTTP 503. It is evidence for graceful failure behavior, not a Gateway outage or successful delivery. The remaining live inbox, Salesforce, CMS publish and FAQ propagation checks were not attempted. Six FAQ freshness cases are prepared separately; no timed update result exists yet.
Evidence handling
Round-one and round-two transcripts, webhook response codes and manifests remain separate. Only round-one raw failures affected by the additional corrections were rerun; passing/partial cases were not repeated. Case transitions use the latest retained result and preserve earlier evidence. Counts are scenario occurrences: W06 appears in two distinct historical test blocks and is not collapsed into one test.
All captures use synthetic identities and local loopback receivers. No real enquiries, shared CMS edits, production/staging settings, deployment or code push occurred.
Final focused retests and source review
Source refresh and final focused fixes
The user requested refreshing the existing source-to-Pinecone pipeline instead of moving runtime lookups to event-management/Sanity APIs. Runtime agenda lookups remain in Pinecone. The new exact read-only eventFacts lookup also reads FAQ/uploads and complete event-page bodies from existing Pinecone namespaces.
The refresh wrote 237 Paris agenda vectors and 145 page vectors covering 33 currently published Paris pages, removing 34 vectors belonging to pages no longer present. Before/current-source/after artifacts are retained in refresh evidence. Other event and knowledge namespaces, website deployment and AI switches were unchanged.
The refreshed live event-management source itself lists seven keynote-format sessions, including Brian Glaser and Stephen Childs, while Eric Mosley is absent from its keynote set. The older local agenda JSON and August six-speaker judge note are stale relative to that source. The earlier suspicion that Pinecone alone had the wrong formats was not supported by the live source check. Both C104 answers matched the current source; their raw failures remain because the original judge note was deliberately preserved. Current source snapshot.
No current volunteer page or volunteer FAQ answer is available. The two C116 replies correctly avoid inventing five team names or linking the removed page, and provide Support instead. Restoring that public content is an editorial decision, not something these tests should force into the answer.
OpenAI R01/R02 passed with complete HR programme, group-ticket and sponsor-package guidance. R07 initially encountered an empty-response runner error; a focused retry passed with manual instructions and Customer Success. OpenAI C78 initially received a blocked judge label because the evaluator lacked policy evidence; the final retest included the captured current FAQ in evaluator notes, unchanged criteria, and a published Attend link, and passed. These evidence changes are recorded in round-four manifest.
Claude H01 initially reached the ticket workflow but repeatedly requested quantity because model extraction could erase saved booking fields. The fix preserves existing booking values through null extraction, recognises an explicit personal “a ticket” request as one, and clears old fields on event/new-booking changes. The final fixed six-message replay passed, with a contact summary, explicit consent and one HTTP 200 local capture. It does not prove real inbox or Salesforce delivery. The transcript’s initial agenda still spans two days despite the visitor’s one-day request; that separate plan-duration issue is recorded rather than hidden by the handoff verdict.
Final H01 transcript · Onsite transcript · Stand plans transcript · Implementation guide.