AI defect fixes — completed failure-only review

Index

Outcome

Reviewed 185 previously non-passing model/block/case occurrences. 177 received real failure/partial-only retests; six confirmed volunteer content occurrences and two separate-session operational checks were retained without paid retries. Across eight adaptive rounds, 301 scenario executions were made; this includes repeat executions of unresolved cases, not that many unique scenarios. Three additional focused quote/information scenario executions and four FAQ-withdrawal response probes are reported separately.

The latest raw grades for this 185-occurrence cohort are 138 pass, 27 partial, 18 fail, 2 blocked. After evidence review: 138 raw passes, 34 obsolete/evaluator-only judgments, 10 content gaps, and 3 operational checks. No verified application defect remains open in this reviewed cohort. Raw non-passes have not been rewritten as passes.

The AI fixes reviewed here are local and not deployed or pushed. OpenAI concierge/agenda uses GPT-5.4-mini; Claude uses Sonnet 4.6. Both use the same OpenAI profile extractor, narrow Jev entry selector, simulator and judge. Jev does not validate responses. No real enquiries were sent; submission URLs were restricted to loopback test receivers. The older billing-blocked full-system job remains stopped.

Original supplied UAT

Only C01–C95, 95 supplied cases per model. Later additions, workflow regressions and FAQ tests are separate. Latest retained grades combine targeted reruns with untouched original results; this is not a new 95-case run.

Model Before pass Latest pass Before fail Latest fail Latest partial Latest blocked
openai 44/95 (46.3%) 82/95 (86.3%) 19 1 11 1
claude 48/95 (50.5%) 81/95 (85.3%) 17 5 8 1

Before and after scores

Before means the original October 5 benchmark, before today’s fixes. After means latest retained results following earlier fixes and this cycle. Full passes only; partial and blocked results remain separate. Evidence-only exclusions use the same denominator before and after, and do not relabel raw verdicts. Content gaps remain scored.

Scope Model Raw before Raw latest Adjusted before Adjusted latest
Original supplied UAT openai 44/95 (46.3%) 82/95 (86.3%) 44/85 (51.8%) 82/85 (96.5%)
Original supplied UAT claude 48/95 (50.5%) 81/95 (85.3%) 48/86 (55.8%) 81/86 (94.2%)
Original supplied UAT Combined raw 92/190 (48.4%) 163/190 (85.8%) — —
Original supplied UAT Combined adjusted 92/171 (53.8%) 163/171 (95.3%) — —
Entire retained benchmark openai 151/279 (54.1%) 239/279 (85.7%) 151/244 (61.9%) 239/244 (98.0%)
Entire retained benchmark claude 168/279 (60.2%) 239/279 (85.7%) 168/247 (68.0%) 239/247 (96.8%)
Entire retained benchmark Combined raw 319/558 (57.2%) 478/558 (85.7%) — —
Entire retained benchmark Combined adjusted 319/491 (65.0%) 478/491 (97.4%) — —

The full benchmark contains 279 occurrences per model (558 combined). 67 occurrences are excluded only from the adjusted denominator: the previous 33 reviewed exclusions plus 34 newly reviewed cases. Exact denominator and exclusion evidence. Model comparisons are descriptive: scenarios were selected adaptively, visitor simulations vary, and content was refreshed. These results do not establish statistical significance.

Benchmark blocks

Each block uses its original cases per model. Latest results retain untouched original cases and replace only corresponding reruns. These are raw grades; content and obsolete expectations are not removed from this table.

Block Model Cases Before pass Latest pass Latest partial Latest fail Latest blocked / error
feedback openai 14 7 11 1 2 0 / 0
feedback claude 14 8 12 0 2 0 / 0
flow-selector openai 23 19 22 0 1 0 / 0
flow-selector claude 23 20 21 1 1 0 / 0
handoff-regression openai 4 2 4 0 0 0 / 0
handoff-regression claude 4 0 3 1 0 0 / 0
knowledge openai 26 8 19 3 4 0 / 0
knowledge claude 26 14 20 2 4 0 / 0
latest-e2e openai 13 13 13 0 0 0 / 0
latest-e2e claude 13 12 13 0 0 0 / 0
original-uat openai 95 44 82 11 1 1 / 0
original-uat claude 95 48 81 8 5 1 / 0
regression openai 38 18 33 1 4 0 / 0
regression claude 38 27 35 0 3 0 / 0
response-policy openai 18 17 18 0 0 0 / 0
response-policy claude 18 18 18 0 0 0 / 0
ticket-precision openai 12 5 7 0 5 0 / 0
ticket-precision claude 12 4 6 1 5 0 / 0
uat-additions openai 24 7 18 2 4 0 / 0
uat-additions claude 24 7 19 2 3 0 / 0
workflow-router openai 12 11 12 0 0 0 / 0
workflow-router claude 12 10 11 1 0 0 / 0

All latest retained raw case results. The original supplied UAT remains its own block; focused replays and FAQ response probes do not inflate these counts.

Missing content

Confirmed content-only gaps affect 10 retained occurrences, not ten site pages. This is a test-coverage percentage, not a measure of missing knowledge-base volume.

Missing area Affected occurrences Evidence / action
Volunteer eligibility, application and team roles C33 Claude; C102 Claude; C116 both; R17 both — 6 No current published page/FAQ supplies the required roles. Publish approved guidance, then sync to the existing Pinecone namespace.
Full sponsor roster names C12 both — 2 Logos mostly lack approved text labels. Add sponsor/company names or alt text; the ingestion fix now includes public referenced names and explicit alt labels. Never infer names from filenames.
Miami speaker destination C21 Claude — 1 Speakers and agenda-schedule pages are unpublished. Publish a destination if intended; otherwise retain honest uncertainty and Support.
Decision-maker percentage C86 Claude — 1 Existing attendance/leadership statistics do not establish a decision-maker percentage. Publish an approved statistic if needed.
Denominator Content-only share
Original supplied UAT, raw 5/190 (2.6%)
Original supplied UAT, adjusted 5/171 (2.9%)
Entire benchmark, raw 10/558 (1.8%)
Entire benchmark, adjusted 10/491 (2.0%)
Reviewed remaining cohort 10/185 (5.4%)

Unpublished detailed startup/free-pass eligibility criteria, attendee parking, venue opening times and some promotional prices are also unavailable. The current tests pass when the bot honestly declines to invent them; that does not make those facts available. All content fixes should follow published page/FAQ → existing Pinecone namespaces, not a different runtime retrieval source.

Application fixes

Relevant code: workflow guards, profile retention, ticket answers, agenda plan, exact event facts, named fact grounding, contact/privacy policy, page ingestion, streaming links, chat endpoint.

Retest rounds and latency

Latency includes profile extraction where that measurement exists. These rounds test different, progressively smaller case sets; medians must not be read as a model speed ranking. Single-chunk deterministic replies are complete text responses, not evidence that model streaming disappeared.

Round Model Cases Pass Partial Fail Median visitor first token, ms Median generation first token, ms Multi-chunk turns / measured turns
1 openai 12 5 2 5 2875 416 10/61
1 claude 11 2 2 7 3291 436 13/54
2 openai 95 43 25 26 3826 1428 119/254
2 claude 75 37 18 20 3532 1019 88/263
3 openai 34 20 10 4 3490 498 16/88
3 claude 27 12 9 6 3474 1015 24/93
4 openai 10 3 7 0 3060 456 3/39
4 claude 14 4 9 1 3483 593 9/45
5 openai 8 4 4 0 2689 521 0/25
5 claude 10 5 4 1 2901 442 3/38
6 openai 1 0 1 0 5655 2842 0/3
6 claude 2 2 0 0 3522 589 2/11
7 openai 1 0 1 0 5246 2967 0/3
8 openai 1 1 0 0 4962 2190 0/4

In the entire latest retained benchmark, captured handoffs with an explicit responseStatus include 106 accepted 2xx and 5 rejected 4xx/5xx. 0 older capture records lack responseStatus and are not independently counted as accepted. This is a latest-per-case count, not the sum of all adaptive attempts. HTTP 503 never counts as a successful submission. Final Claude forced-failure evidence.

The final OpenAI startup replay uses the actual earlier failing visitor messages, including Award details, price and closing address. This avoids claiming a different successful simulator conversation necessarily reproduced the original defect. Replay and criteria record.

Focused Miami quote recovery

The random C49 simulator passed a sponsorship route in round six, which did not verify the earlier failed Miami ticket branch. We therefore replayed the actual failed quote intent and turn order, using explicit synthetic UAT name/email/company/role fields. The first replay caught one further consent loop: “send this group quote enquiry with my consent” repeated the summary. That phrase is now accepted only after a handoff question; conditions, cancellation and field changes remain rejected.

The second replay passes: four attendees are retained; no pass is invented; the complete summary and explicit confirmation precede exactly one rejected 503 attempt with Delegate UNA intent. The response honestly reports failure and Support, then closes without resuming collection. A separate one-turn replay of the earlier polite pricing question also passes, with no company/role/contact questions and no webhook. These are three additional focused scenario executions (one fail, two passes), kept separate from the 301 cohort executions and the 558 retained benchmark scores.

Earlier failed quote replay · Passing exact synthetic quote replay · Passing read-only pricing turn.

FAQ freshness and withdrawal

The separate live Pinecone freshness suite passed seven mechanical scenarios: add, edit, refresh in an existing chat, shortened-file stale-tail removal, FAQ withdrawal, last-FAQ deletion and file deletion. Observed answer update samples were roughly 9–11 seconds, with one existing-chat sample at 2.87 seconds. These are measurements from this run, not a Sanity publish-to-answer SLA. Original freshness evidence.

Mechanical deletion success initially hid two response defects: the bot could invent a removed room detail, and quote a removed file’s welcome phrase. Prompt-only grounding was insufficient in the first response retest, which is retained as failed evidence. The conservative same-entity/same-field source check now prevents unsupported answers. The final response tests are 4/4 verified across OpenAI and Claude: both removed FAQ and removed file queries return honest inability to confirm with Support, not stale/fabricated details. Response times were 899–1,393 ms in these probes; no model token stream is needed for the deterministic fallback.

Only unique, empty test namespaces were modified. Removal was confirmed before querying the app. Cleanup verified empty namespaces in two consecutive observations; Claude’s intermittent stale list results took 17 seconds to settle. Temporary environment copies were removed and no enquiry was captured. OpenAI responses · Claude responses · OpenAI cleanup · Claude cleanup · Earlier unsuccessful attempt.

Remaining raw non-passes

These are the non-passing grades in the 185-occurrence cohort. The other earlier reviewed exclusions remain itemised in retained scores. “Reviewed” is a finding, not a replacement judge verdict.

Model Block / case Raw Reviewed cause Evidence
openai feedback / F06 partial evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. Transcript
openai flow-selector / REGRESSION-R06 fail evaluator_or_obsolete_expectation: Generic floorplan questions use Support under the approved policy. No published floorplan was found; claiming it is categorically non-public is unsupported. Clear stand-builder/exhibitor queries separately use Customer Success. Transcript
openai knowledge / K06 fail evaluator_or_obsolete_expectation: Ticket transfers use Support under the approved contact policy. The current answer preserves availability and Terms & Conditions. Transcript
openai knowledge / K07 fail evaluator_or_obsolete_expectation: Ticket upgrades use Support under the approved contact policy; availability is conditional and no fee was invented. Transcript
openai knowledge / K14 partial evaluator_or_obsolete_expectation: Invoice/payment enquiries are general ticket support; the former Customer Success requirement is obsolete. Transcript
openai original-uat / C01 partial evaluator_or_obsolete_expectation: The four required visitor fields, reviewed summary and final confirmation precede an accepted 200 handoff. Mandatory company size, seniority and automatic free HR Leader qualification are obsolete/unmaintained requirements. The literal request to tell the visitor about their company is subject to the current personal-data policy. Transcript
openai original-uat / C04 partial evaluator_or_obsolete_expectation: The response matches the captured structured prices: EUR 3,195 + VAT General Attendee; EUR 5,495 + VAT Diamond; Explorer price on application. Evaluator lack of source access is not evidence of invention. Transcript · Source
openai original-uat / C05 partial evaluator_or_obsolete_expectation: This is an informational group-ticket question. Requiring lead qualification and consent before a requested handoff conflicts with the approved contract. Transcript
openai original-uat / C06 partial evaluator_or_obsolete_expectation: Attend and the published General Attendee page are valid registration destinations. Requiring an additional external checkout URL is not in the contract. All required fields, summary and consent precede exactly one accepted 200 handoff. Transcript
openai original-uat / C10 partial evaluator_or_obsolete_expectation: Current contract reuses company/role and requires name, work email, company and role, with one final summary confirmation. Mandatory discrete company size/location/industry fields and their collection order are obsolete. Accepted 200 evidence is retained. Transcript
openai original-uat / C12 partial content_gap: Sponsor names are largely image-only assets without approved text labels. Named public profiles and explicit asset alt text now ingest; filenames are not trusted sponsor names. The full sponsor roster still needs editorial labels. Transcript · Source
openai original-uat / C26 partial evaluator_or_obsolete_expectation: The requested report format is preserved with real published destinations. The visitor did not require recent reports; an evaluator preference for newer reports is not a failed freshness contract. Transcript
openai original-uat / C42 partial evaluator_or_obsolete_expectation: A generic attendee-list request is refused without disclosure. Support is correct in the absence of clear sponsor/exhibitor context; the test assumed a commercial context. Transcript
openai original-uat / C47 blocked operational_unverified: Separate browser/session identity isolation was not exercised by this single-chat simulator. Transcript
openai original-uat / C50 partial operational_unverified: Malformed email correction, accurate summary and accepted 200 handoff work. Actual downstream SMTP hard-bounce handling is outside the local receiver and remains unverified. Transcript
openai original-uat / C53 partial evaluator_or_obsolete_expectation: Employer-side HR practitioner guidance and application review are preserved. Specific eligibility thresholds are not published; mandatory lead qualification for information is obsolete. Transcript
openai original-uat / C61 partial evaluator_or_obsolete_expectation: The visitor explicitly requested only an exhibition stand. The published sponsorship/exhibition destination, four fields, complete summary and consent precede one accepted 200 handoff. Adding ticket choice is an irrelevant simulator instruction. Transcript
openai regression / R05 fail evaluator_or_obsolete_expectation: The current approved general contact is Support; the former additional events@unleash.ai requirement is obsolete. Transcript
openai regression / R06 fail evaluator_or_obsolete_expectation: Generic floorplan Support routing is approved; no unsupported city map or claim of public availability is made. A mandatory Customer Success route presupposes absent SPEX context. Transcript
openai regression / R09 fail evaluator_or_obsolete_expectation: Only Support is required for this general contact question under the current policy; the additional events address is obsolete. Transcript
openai regression / R17 fail content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source
openai regression / R29 partial evaluator_or_obsolete_expectation: The current event dates are correct. The user asks for dates; an extra weekday display is an evaluator formatting preference, not an incorrect date. Transcript
openai ticket-precision / T02 fail evaluator_or_obsolete_expectation: The response uses the current structured Diamond price, Explorer application status and a concrete preference question. Only the old EUR 4,995 price expectation remains invalid. Transcript · Source
openai ticket-precision / T05 fail evaluator_or_obsolete_expectation: The unit price EUR 3,195 + VAT is verified in the captured structured source. The expected EUR 2,995 and its total are obsolete. Transcript · Source
openai uat-additions / C105 partial evaluator_or_obsolete_expectation: Amanda Poole is listed in the current Workhuman Forum keynote. The full current seven-speaker list matches the source; the historical six-speaker fixture is stale. Transcript · Source
openai uat-additions / C111 partial evaluator_or_obsolete_expectation: Accepted ticket handoff with all four required fields, accurate summary and final consent. Name-first ordering is not required by the current conversational contract. Transcript
openai uat-additions / C116 fail content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source
claude flow-selector / REGRESSION-R06 fail evaluator_or_obsolete_expectation: Generic floorplan questions use Support under the approved policy. No published floorplan was found; claiming it is categorically non-public is unsupported. Clear stand-builder/exhibitor queries separately use Customer Success. Transcript
claude handoff-regression / F06 partial evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. Transcript
claude original-uat / C01 partial evaluator_or_obsolete_expectation: The four required visitor fields, reviewed summary and final confirmation precede an accepted 200 handoff. Mandatory company size, seniority and automatic free HR Leader qualification are obsolete/unmaintained requirements. The literal request to tell the visitor about their company is subject to the current personal-data policy. Transcript
claude original-uat / C04 partial evaluator_or_obsolete_expectation: The response matches the captured structured prices: EUR 3,195 + VAT General Attendee; EUR 5,495 + VAT Diamond; Explorer price on application. Evaluator lack of source access is not evidence of invention. Transcript · Source
claude original-uat / C06 partial evaluator_or_obsolete_expectation: Attend and the published General Attendee page are valid registration destinations. Requiring an additional external checkout URL is not in the contract. All required fields, summary and consent precede exactly one accepted 200 handoff. Transcript
claude original-uat / C10 partial evaluator_or_obsolete_expectation: Current contract reuses company/role and requires name, work email, company and role, with one final summary confirmation. Mandatory discrete company size/location/industry fields and their collection order are obsolete. Accepted 200 evidence is retained. Transcript
claude original-uat / C12 partial content_gap: Sponsor names are largely image-only assets without approved text labels. Named public profiles and explicit asset alt text now ingest; filenames are not trusted sponsor names. The full sponsor roster still needs editorial labels. Transcript · Source
claude original-uat / C21 partial content_gap: Miami speakers and agenda-schedule pages are unpublished; the honest uncertainty/contact answer is correct, but the expected public destination cannot be provided. Transcript · Source
claude original-uat / C33 fail content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source
claude original-uat / C47 blocked operational_unverified: Separate browser/session identity isolation was not exercised by this single-chat simulator. Transcript
claude original-uat / C86 partial content_gap: The current published FAQ has attendee/seniority context but no verified decision-maker percentage. That percentage cannot be invented. Transcript · Source
claude regression / R06 fail evaluator_or_obsolete_expectation: Generic floorplan Support routing is approved; no unsupported city map or claim of public availability is made. A mandatory Customer Success route presupposes absent SPEX context. Transcript
claude regression / R09 fail evaluator_or_obsolete_expectation: Only Support is required for this general contact question under the current policy; the additional events address is obsolete. Transcript
claude regression / R17 fail content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source
claude ticket-precision / F06 partial evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. Transcript
claude ticket-precision / T02 fail evaluator_or_obsolete_expectation: The response uses the current structured Diamond price, Explorer application status and a concrete preference question. Only the old EUR 4,995 price expectation remains invalid. Transcript · Source
claude ticket-precision / T05 fail evaluator_or_obsolete_expectation: The unit price EUR 3,195 + VAT is verified in the captured structured source. The expected EUR 2,995 and its total are obsolete. Transcript · Source
claude uat-additions / C102 partial content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source
claude uat-additions / C114 partial evaluator_or_obsolete_expectation: Known company is reused, quantity is asked before it has been supplied, and final summary/consent precede one accepted 200 handoff. An information-only turn need not end with qualification or a post-send repeat summary. Transcript
claude uat-additions / C116 fail content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. Transcript · Source

Verification and operational limits

Documentation and change log

The complete October fix record documents each observed problem, resulting behavior, code/regression links and exact current code extracts. It explains routing and known-fact memory, natural consent, group quotes and quantities, day-scoped agendas, source-only startup/Award replies, withdrawn-fact guards, source ingestion, contact/privacy policy and streaming link delivery.

The technical architecture, simple call guide, tool guide, facts guide, privacy/submissions guide, editor operations and launch controls have been updated. Earlier unresolved agenda/profile/confirmation statements are superseded by the final evidence; historical test artifacts remain unchanged.

Deployment status: these AI fixes remain local and unpushed. The scoped published-page Pinecone refresh was applied; no Sanity documents or published launch/model settings were edited during this fix cycle. Final test apps were stopped, test namespaces were verified empty, and user dev servers were untouched. No real enquiries were sent.

The separate earlier event-route isolation repair was deployed to main at 6ba30a04: saved production smoke shows the unpublished Miami agenda paths returning 404 while Paris routes and event roots return 200. Its 11 isolated regressions are separate from the AI counts. Deployment evidence · Saved production smoke. No deployment was made during this AI fix/documentation cycle.

Report rebuild and artifact contracts · Final local verification · Cleanup. Rebuilding this report only reads retained evidence; it does not restart paid tests.

Evidence

All 185 reviewed occurrences · JSON review ledger · Latest retained 558 raw results · Scores/exclusion denominator · Per-round execution/latency measurements · Frozen final source manifest · Source refresh · Published/raw ingestion evidence · Miami publication check.