AI defect fixes — completed failure-only review
Index
- Outcome
- Original supplied UAT
- Before and after scores
- Benchmark blocks
- Missing content
- Application fixes
- Retest rounds and latency
- Focused Miami quote recovery
- FAQ freshness and withdrawal
- Remaining raw non-passes
- Verification and operational limits
- Documentation and change log
- Evidence
Outcome
Reviewed 185 previously non-passing model/block/case occurrences. 177 received real failure/partial-only retests; six confirmed volunteer content occurrences and two separate-session operational checks were retained without paid retries. Across eight adaptive rounds, 301 scenario executions were made; this includes repeat executions of unresolved cases, not that many unique scenarios. Three additional focused quote/information scenario executions and four FAQ-withdrawal response probes are reported separately.
The latest raw grades for this 185-occurrence cohort are 138 pass, 27 partial, 18 fail, 2 blocked. After evidence review: 138 raw passes, 34 obsolete/evaluator-only judgments, 10 content gaps, and 3 operational checks. No verified application defect remains open in this reviewed cohort. Raw non-passes have not been rewritten as passes.
The AI fixes reviewed here are local and not deployed or pushed. OpenAI concierge/agenda uses GPT-5.4-mini; Claude uses Sonnet 4.6. Both use the same OpenAI profile extractor, narrow Jev entry selector, simulator and judge. Jev does not validate responses. No real enquiries were sent; submission URLs were restricted to loopback test receivers. The older billing-blocked full-system job remains stopped.
Original supplied UAT
Only C01–C95, 95 supplied cases per model. Later additions, workflow regressions and FAQ tests are separate. Latest retained grades combine targeted reruns with untouched original results; this is not a new 95-case run.
| Model | Before pass | Latest pass | Before fail | Latest fail | Latest partial | Latest blocked |
|---|---|---|---|---|---|---|
| openai | 44/95 (46.3%) | 82/95 (86.3%) | 19 | 1 | 11 | 1 |
| claude | 48/95 (50.5%) | 81/95 (85.3%) | 17 | 5 | 8 | 1 |
Before and after scores
Before means the original October 5 benchmark, before today’s fixes. After means latest retained results following earlier fixes and this cycle. Full passes only; partial and blocked results remain separate. Evidence-only exclusions use the same denominator before and after, and do not relabel raw verdicts. Content gaps remain scored.
| Scope | Model | Raw before | Raw latest | Adjusted before | Adjusted latest |
|---|---|---|---|---|---|
| Original supplied UAT | openai | 44/95 (46.3%) | 82/95 (86.3%) | 44/85 (51.8%) | 82/85 (96.5%) |
| Original supplied UAT | claude | 48/95 (50.5%) | 81/95 (85.3%) | 48/86 (55.8%) | 81/86 (94.2%) |
| Original supplied UAT | Combined raw | 92/190 (48.4%) | 163/190 (85.8%) | — | — |
| Original supplied UAT | Combined adjusted | 92/171 (53.8%) | 163/171 (95.3%) | — | — |
| Entire retained benchmark | openai | 151/279 (54.1%) | 239/279 (85.7%) | 151/244 (61.9%) | 239/244 (98.0%) |
| Entire retained benchmark | claude | 168/279 (60.2%) | 239/279 (85.7%) | 168/247 (68.0%) | 239/247 (96.8%) |
| Entire retained benchmark | Combined raw | 319/558 (57.2%) | 478/558 (85.7%) | — | — |
| Entire retained benchmark | Combined adjusted | 319/491 (65.0%) | 478/491 (97.4%) | — | — |
The full benchmark contains 279 occurrences per model (558 combined). 67 occurrences are excluded only from the adjusted denominator: the previous 33 reviewed exclusions plus 34 newly reviewed cases. Exact denominator and exclusion evidence. Model comparisons are descriptive: scenarios were selected adaptively, visitor simulations vary, and content was refreshed. These results do not establish statistical significance.
Benchmark blocks
Each block uses its original cases per model. Latest results retain untouched original cases and replace only corresponding reruns. These are raw grades; content and obsolete expectations are not removed from this table.
| Block | Model | Cases | Before pass | Latest pass | Latest partial | Latest fail | Latest blocked / error |
|---|---|---|---|---|---|---|---|
| feedback | openai | 14 | 7 | 11 | 1 | 2 | 0 / 0 |
| feedback | claude | 14 | 8 | 12 | 0 | 2 | 0 / 0 |
| flow-selector | openai | 23 | 19 | 22 | 0 | 1 | 0 / 0 |
| flow-selector | claude | 23 | 20 | 21 | 1 | 1 | 0 / 0 |
| handoff-regression | openai | 4 | 2 | 4 | 0 | 0 | 0 / 0 |
| handoff-regression | claude | 4 | 0 | 3 | 1 | 0 | 0 / 0 |
| knowledge | openai | 26 | 8 | 19 | 3 | 4 | 0 / 0 |
| knowledge | claude | 26 | 14 | 20 | 2 | 4 | 0 / 0 |
| latest-e2e | openai | 13 | 13 | 13 | 0 | 0 | 0 / 0 |
| latest-e2e | claude | 13 | 12 | 13 | 0 | 0 | 0 / 0 |
| original-uat | openai | 95 | 44 | 82 | 11 | 1 | 1 / 0 |
| original-uat | claude | 95 | 48 | 81 | 8 | 5 | 1 / 0 |
| regression | openai | 38 | 18 | 33 | 1 | 4 | 0 / 0 |
| regression | claude | 38 | 27 | 35 | 0 | 3 | 0 / 0 |
| response-policy | openai | 18 | 17 | 18 | 0 | 0 | 0 / 0 |
| response-policy | claude | 18 | 18 | 18 | 0 | 0 | 0 / 0 |
| ticket-precision | openai | 12 | 5 | 7 | 0 | 5 | 0 / 0 |
| ticket-precision | claude | 12 | 4 | 6 | 1 | 5 | 0 / 0 |
| uat-additions | openai | 24 | 7 | 18 | 2 | 4 | 0 / 0 |
| uat-additions | claude | 24 | 7 | 19 | 2 | 3 | 0 / 0 |
| workflow-router | openai | 12 | 11 | 12 | 0 | 0 | 0 / 0 |
| workflow-router | claude | 12 | 10 | 11 | 1 | 0 | 0 / 0 |
All latest retained raw case results. The original supplied UAT remains its own block; focused replays and FAQ response probes do not inflate these counts.
Missing content
Confirmed content-only gaps affect 10 retained occurrences, not ten site pages. This is a test-coverage percentage, not a measure of missing knowledge-base volume.
| Missing area | Affected occurrences | Evidence / action |
|---|---|---|
| Volunteer eligibility, application and team roles | C33 Claude; C102 Claude; C116 both; R17 both — 6 | No current published page/FAQ supplies the required roles. Publish approved guidance, then sync to the existing Pinecone namespace. |
| Full sponsor roster names | C12 both — 2 | Logos mostly lack approved text labels. Add sponsor/company names or alt text; the ingestion fix now includes public referenced names and explicit alt labels. Never infer names from filenames. |
| Miami speaker destination | C21 Claude — 1 | Speakers and agenda-schedule pages are unpublished. Publish a destination if intended; otherwise retain honest uncertainty and Support. |
| Decision-maker percentage | C86 Claude — 1 | Existing attendance/leadership statistics do not establish a decision-maker percentage. Publish an approved statistic if needed. |
| Denominator | Content-only share |
|---|---|
| Original supplied UAT, raw | 5/190 (2.6%) |
| Original supplied UAT, adjusted | 5/171 (2.9%) |
| Entire benchmark, raw | 10/558 (1.8%) |
| Entire benchmark, adjusted | 10/491 (2.0%) |
| Reviewed remaining cohort | 10/185 (5.4%) |
Unpublished detailed startup/free-pass eligibility criteria, attendee parking, venue opening times and some promotional prices are also unavailable. The current tests pass when the bot honestly declines to invent them; that does not make those facts available. All content fixes should follow published page/FAQ → existing Pinecone namespaces, not a different runtime retrieval source.
Application fixes
- Workflow and consent: active workflows retain their state; cancellation clears handoff state; natural confirmation resumes the intended send; saved consent still requires reviewing the current request. SPEX summaries include contact fields and the visitor’s objective before submission.
- Tickets: current structured prices and Explorer application status are preserved; quantity ranges remain estimates; repeated choice-help advances to a recommendation. A requested group quote can proceed without forcing a pass choice, including Miami when verified prices are unavailable.
- Agenda: requested days, source year, non-overlap and access restrictions are enforced. Missing interests are asked once; corrected recipients and saved profile fields survive. Browsing lists are validated into a plan before email handoff.
- Information and grounding: logistical and informational questions do not initiate lead capture. Stage and speaker queries stay scoped. Media format requests retain reports/webinars/research; keynote rosters use current source evidence. Removed named FAQ facts require supporting evidence for the same entity and field.
- Startup follow-ups: eligibility, package pricing, badge registration and the separate Award application get distinct answers. Unsupported approval/price claims are withheld, and acknowledgements close cleanly. The Award reply copies the published competition description/application URL instead of sending users to badge checkout.
- Contact and privacy: general/ticket/refund/privacy/failure uses Support; clear sponsor/exhibitor/programme questions use Customer Success. Volunteering contact details is not a personal-data export. Mentioning “that email address” is not a new generic contact question.
- Source ingestion: public referenced company/profile names and explicit image alt text are extracted. Incremental reference changes refresh related pages. Published/raw scope equivalence was checked on real data; five published Paris pages were refreshed into 22 vectors. No Sanity documents were edited.
- Links and delivery: unsafe/empty placeholder links become plain text; quoted URLs are normalized; step boundaries retain readable paragraphs. Link normalization buffers only the current link so ordinary model text continues streaming. Accepted/rejected handoff tool outcomes own the final delivery response.
Relevant code: workflow guards, profile retention, ticket answers, agenda plan, exact event facts, named fact grounding, contact/privacy policy, page ingestion, streaming links, chat endpoint.
Retest rounds and latency
Latency includes profile extraction where that measurement exists. These rounds test different, progressively smaller case sets; medians must not be read as a model speed ranking. Single-chunk deterministic replies are complete text responses, not evidence that model streaming disappeared.
| Round | Model | Cases | Pass | Partial | Fail | Median visitor first token, ms | Median generation first token, ms | Multi-chunk turns / measured turns |
|---|---|---|---|---|---|---|---|---|
| 1 | openai | 12 | 5 | 2 | 5 | 2875 | 416 | 10/61 |
| 1 | claude | 11 | 2 | 2 | 7 | 3291 | 436 | 13/54 |
| 2 | openai | 95 | 43 | 25 | 26 | 3826 | 1428 | 119/254 |
| 2 | claude | 75 | 37 | 18 | 20 | 3532 | 1019 | 88/263 |
| 3 | openai | 34 | 20 | 10 | 4 | 3490 | 498 | 16/88 |
| 3 | claude | 27 | 12 | 9 | 6 | 3474 | 1015 | 24/93 |
| 4 | openai | 10 | 3 | 7 | 0 | 3060 | 456 | 3/39 |
| 4 | claude | 14 | 4 | 9 | 1 | 3483 | 593 | 9/45 |
| 5 | openai | 8 | 4 | 4 | 0 | 2689 | 521 | 0/25 |
| 5 | claude | 10 | 5 | 4 | 1 | 2901 | 442 | 3/38 |
| 6 | openai | 1 | 0 | 1 | 0 | 5655 | 2842 | 0/3 |
| 6 | claude | 2 | 2 | 0 | 0 | 3522 | 589 | 2/11 |
| 7 | openai | 1 | 0 | 1 | 0 | 5246 | 2967 | 0/3 |
| 8 | openai | 1 | 1 | 0 | 0 | 4962 | 2190 | 0/4 |
In the entire latest retained benchmark, captured handoffs with an explicit responseStatus include 106 accepted 2xx and 5 rejected 4xx/5xx. 0 older capture records lack responseStatus and are not independently counted as accepted. This is a latest-per-case count, not the sum of all adaptive attempts. HTTP 503 never counts as a successful submission. Final Claude forced-failure evidence.
The final OpenAI startup replay uses the actual earlier failing visitor messages, including Award details, price and closing address. This avoids claiming a different successful simulator conversation necessarily reproduced the original defect. Replay and criteria record.
Focused Miami quote recovery
The random C49 simulator passed a sponsorship route in round six, which did not verify the earlier failed Miami ticket branch. We therefore replayed the actual failed quote intent and turn order, using explicit synthetic UAT name/email/company/role fields. The first replay caught one further consent loop: “send this group quote enquiry with my consent” repeated the summary. That phrase is now accepted only after a handoff question; conditions, cancellation and field changes remain rejected.
The second replay passes: four attendees are retained; no pass is invented; the complete summary and explicit confirmation precede exactly one rejected 503 attempt with Delegate UNA intent. The response honestly reports failure and Support, then closes without resuming collection. A separate one-turn replay of the earlier polite pricing question also passes, with no company/role/contact questions and no webhook. These are three additional focused scenario executions (one fail, two passes), kept separate from the 301 cohort executions and the 558 retained benchmark scores.
Earlier failed quote replay · Passing exact synthetic quote replay · Passing read-only pricing turn.
FAQ freshness and withdrawal
The separate live Pinecone freshness suite passed seven mechanical scenarios: add, edit, refresh in an existing chat, shortened-file stale-tail removal, FAQ withdrawal, last-FAQ deletion and file deletion. Observed answer update samples were roughly 9–11 seconds, with one existing-chat sample at 2.87 seconds. These are measurements from this run, not a Sanity publish-to-answer SLA. Original freshness evidence.
Mechanical deletion success initially hid two response defects: the bot could invent a removed room detail, and quote a removed file’s welcome phrase. Prompt-only grounding was insufficient in the first response retest, which is retained as failed evidence. The conservative same-entity/same-field source check now prevents unsupported answers. The final response tests are 4/4 verified across OpenAI and Claude: both removed FAQ and removed file queries return honest inability to confirm with Support, not stale/fabricated details. Response times were 899–1,393 ms in these probes; no model token stream is needed for the deterministic fallback.
Only unique, empty test namespaces were modified. Removal was confirmed before querying the app. Cleanup verified empty namespaces in two consecutive observations; Claude’s intermittent stale list results took 17 seconds to settle. Temporary environment copies were removed and no enquiry was captured. OpenAI responses · Claude responses · OpenAI cleanup · Claude cleanup · Earlier unsuccessful attempt.
Remaining raw non-passes
These are the non-passing grades in the 185-occurrence cohort. The other earlier reviewed exclusions remain itemised in retained scores. “Reviewed” is a finding, not a replacement judge verdict.
| Model | Block / case | Raw | Reviewed cause | Evidence |
|---|---|---|---|---|
| openai | feedback / F06 | partial | evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. | Transcript |
| openai | flow-selector / REGRESSION-R06 | fail | evaluator_or_obsolete_expectation: Generic floorplan questions use Support under the approved policy. No published floorplan was found; claiming it is categorically non-public is unsupported. Clear stand-builder/exhibitor queries separately use Customer Success. | Transcript |
| openai | knowledge / K06 | fail | evaluator_or_obsolete_expectation: Ticket transfers use Support under the approved contact policy. The current answer preserves availability and Terms & Conditions. | Transcript |
| openai | knowledge / K07 | fail | evaluator_or_obsolete_expectation: Ticket upgrades use Support under the approved contact policy; availability is conditional and no fee was invented. | Transcript |
| openai | knowledge / K14 | partial | evaluator_or_obsolete_expectation: Invoice/payment enquiries are general ticket support; the former Customer Success requirement is obsolete. | Transcript |
| openai | original-uat / C01 | partial | evaluator_or_obsolete_expectation: The four required visitor fields, reviewed summary and final confirmation precede an accepted 200 handoff. Mandatory company size, seniority and automatic free HR Leader qualification are obsolete/unmaintained requirements. The literal request to tell the visitor about their company is subject to the current personal-data policy. | Transcript |
| openai | original-uat / C04 | partial | evaluator_or_obsolete_expectation: The response matches the captured structured prices: EUR 3,195 + VAT General Attendee; EUR 5,495 + VAT Diamond; Explorer price on application. Evaluator lack of source access is not evidence of invention. | Transcript · Source |
| openai | original-uat / C05 | partial | evaluator_or_obsolete_expectation: This is an informational group-ticket question. Requiring lead qualification and consent before a requested handoff conflicts with the approved contract. | Transcript |
| openai | original-uat / C06 | partial | evaluator_or_obsolete_expectation: Attend and the published General Attendee page are valid registration destinations. Requiring an additional external checkout URL is not in the contract. All required fields, summary and consent precede exactly one accepted 200 handoff. | Transcript |
| openai | original-uat / C10 | partial | evaluator_or_obsolete_expectation: Current contract reuses company/role and requires name, work email, company and role, with one final summary confirmation. Mandatory discrete company size/location/industry fields and their collection order are obsolete. Accepted 200 evidence is retained. | Transcript |
| openai | original-uat / C12 | partial | content_gap: Sponsor names are largely image-only assets without approved text labels. Named public profiles and explicit asset alt text now ingest; filenames are not trusted sponsor names. The full sponsor roster still needs editorial labels. | Transcript · Source |
| openai | original-uat / C26 | partial | evaluator_or_obsolete_expectation: The requested report format is preserved with real published destinations. The visitor did not require recent reports; an evaluator preference for newer reports is not a failed freshness contract. | Transcript |
| openai | original-uat / C42 | partial | evaluator_or_obsolete_expectation: A generic attendee-list request is refused without disclosure. Support is correct in the absence of clear sponsor/exhibitor context; the test assumed a commercial context. | Transcript |
| openai | original-uat / C47 | blocked | operational_unverified: Separate browser/session identity isolation was not exercised by this single-chat simulator. | Transcript |
| openai | original-uat / C50 | partial | operational_unverified: Malformed email correction, accurate summary and accepted 200 handoff work. Actual downstream SMTP hard-bounce handling is outside the local receiver and remains unverified. | Transcript |
| openai | original-uat / C53 | partial | evaluator_or_obsolete_expectation: Employer-side HR practitioner guidance and application review are preserved. Specific eligibility thresholds are not published; mandatory lead qualification for information is obsolete. | Transcript |
| openai | original-uat / C61 | partial | evaluator_or_obsolete_expectation: The visitor explicitly requested only an exhibition stand. The published sponsorship/exhibition destination, four fields, complete summary and consent precede one accepted 200 handoff. Adding ticket choice is an irrelevant simulator instruction. | Transcript |
| openai | regression / R05 | fail | evaluator_or_obsolete_expectation: The current approved general contact is Support; the former additional events@unleash.ai requirement is obsolete. | Transcript |
| openai | regression / R06 | fail | evaluator_or_obsolete_expectation: Generic floorplan Support routing is approved; no unsupported city map or claim of public availability is made. A mandatory Customer Success route presupposes absent SPEX context. | Transcript |
| openai | regression / R09 | fail | evaluator_or_obsolete_expectation: Only Support is required for this general contact question under the current policy; the additional events address is obsolete. | Transcript |
| openai | regression / R17 | fail | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
| openai | regression / R29 | partial | evaluator_or_obsolete_expectation: The current event dates are correct. The user asks for dates; an extra weekday display is an evaluator formatting preference, not an incorrect date. | Transcript |
| openai | ticket-precision / T02 | fail | evaluator_or_obsolete_expectation: The response uses the current structured Diamond price, Explorer application status and a concrete preference question. Only the old EUR 4,995 price expectation remains invalid. | Transcript · Source |
| openai | ticket-precision / T05 | fail | evaluator_or_obsolete_expectation: The unit price EUR 3,195 + VAT is verified in the captured structured source. The expected EUR 2,995 and its total are obsolete. | Transcript · Source |
| openai | uat-additions / C105 | partial | evaluator_or_obsolete_expectation: Amanda Poole is listed in the current Workhuman Forum keynote. The full current seven-speaker list matches the source; the historical six-speaker fixture is stale. | Transcript · Source |
| openai | uat-additions / C111 | partial | evaluator_or_obsolete_expectation: Accepted ticket handoff with all four required fields, accurate summary and final consent. Name-first ordering is not required by the current conversational contract. | Transcript |
| openai | uat-additions / C116 | fail | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
| claude | flow-selector / REGRESSION-R06 | fail | evaluator_or_obsolete_expectation: Generic floorplan questions use Support under the approved policy. No published floorplan was found; claiming it is categorically non-public is unsupported. Clear stand-builder/exhibitor queries separately use Customer Success. | Transcript |
| claude | handoff-regression / F06 | partial | evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. | Transcript |
| claude | original-uat / C01 | partial | evaluator_or_obsolete_expectation: The four required visitor fields, reviewed summary and final confirmation precede an accepted 200 handoff. Mandatory company size, seniority and automatic free HR Leader qualification are obsolete/unmaintained requirements. The literal request to tell the visitor about their company is subject to the current personal-data policy. | Transcript |
| claude | original-uat / C04 | partial | evaluator_or_obsolete_expectation: The response matches the captured structured prices: EUR 3,195 + VAT General Attendee; EUR 5,495 + VAT Diamond; Explorer price on application. Evaluator lack of source access is not evidence of invention. | Transcript · Source |
| claude | original-uat / C06 | partial | evaluator_or_obsolete_expectation: Attend and the published General Attendee page are valid registration destinations. Requiring an additional external checkout URL is not in the contract. All required fields, summary and consent precede exactly one accepted 200 handoff. | Transcript |
| claude | original-uat / C10 | partial | evaluator_or_obsolete_expectation: Current contract reuses company/role and requires name, work email, company and role, with one final summary confirmation. Mandatory discrete company size/location/industry fields and their collection order are obsolete. Accepted 200 evidence is retained. | Transcript |
| claude | original-uat / C12 | partial | content_gap: Sponsor names are largely image-only assets without approved text labels. Named public profiles and explicit asset alt text now ingest; filenames are not trusted sponsor names. The full sponsor roster still needs editorial labels. | Transcript · Source |
| claude | original-uat / C21 | partial | content_gap: Miami speakers and agenda-schedule pages are unpublished; the honest uncertainty/contact answer is correct, but the expected public destination cannot be provided. | Transcript · Source |
| claude | original-uat / C33 | fail | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
| claude | original-uat / C47 | blocked | operational_unverified: Separate browser/session identity isolation was not exercised by this single-chat simulator. | Transcript |
| claude | original-uat / C86 | partial | content_gap: The current published FAQ has attendee/seniority context but no verified decision-maker percentage. That percentage cannot be invented. | Transcript · Source |
| claude | regression / R06 | fail | evaluator_or_obsolete_expectation: Generic floorplan Support routing is approved; no unsupported city map or claim of public availability is made. A mandatory Customer Success route presupposes absent SPEX context. | Transcript |
| claude | regression / R09 | fail | evaluator_or_obsolete_expectation: Only Support is required for this general contact question under the current policy; the additional events address is obsolete. | Transcript |
| claude | regression / R17 | fail | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
| claude | ticket-precision / F06 | partial | evaluator_or_obsolete_expectation: The transcript compares current passes, retains quantity and known fields, then advances to a recommendation and Day 1 choice. The judge penalises an initial question asked before any help request while acknowledging that it is acceptable. | Transcript |
| claude | ticket-precision / T02 | fail | evaluator_or_obsolete_expectation: The response uses the current structured Diamond price, Explorer application status and a concrete preference question. Only the old EUR 4,995 price expectation remains invalid. | Transcript · Source |
| claude | ticket-precision / T05 | fail | evaluator_or_obsolete_expectation: The unit price EUR 3,195 + VAT is verified in the captured structured source. The expected EUR 2,995 and its total are obsolete. | Transcript · Source |
| claude | uat-additions / C102 | partial | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
| claude | uat-additions / C114 | partial | evaluator_or_obsolete_expectation: Known company is reused, quantity is asked before it has been supplied, and final summary/consent precede one accepted 200 handoff. An information-only turn need not end with qualification or a post-send repeat summary. | Transcript |
| claude | uat-additions / C116 | fail | content_gap: Current volunteering roles, team breakdown and application/eligibility details are absent from the reviewed published pages and FAQs. Safe Support fallback is not a complete content answer. | Transcript · Source |
Verification and operational limits
- Latest web library verification: 264 tests pass, 0 fail. Shared AI package: 39 pass, 0 fail. Focused regressions cover quantity ranges, corrected email/consent, choice-help, source-only Award details, closing messages, safe links, and withdrawn FAQ facts.
- Touched-file Biome checks pass. The web typecheck has existing failures elsewhere; the captured final scan contains no errors in touched source files. It is not a clean whole-repository typecheck.
- Real GROQ ingestion changes were checked in both published and raw perspectives; document scope stayed unchanged.
- Not verified: actual Zapier→Salesforce/Postmark delivery, SMTP bounce handling, browser/new-session identity isolation in C47, live deployment/cookie/production visibility, configured Sanity publish-webhook filters, PDF extraction, and sustained publish-to-answer latency. The local test receiver cannot certify these.
- Source data, adaptive visitor conversations and judge outcomes can change. Zero reviewed open application defects in this cohort is not a promise that every possible conversation succeeds.
Documentation and change log
The complete October fix record documents each observed problem, resulting behavior, code/regression links and exact current code extracts. It explains routing and known-fact memory, natural consent, group quotes and quantities, day-scoped agendas, source-only startup/Award replies, withdrawn-fact guards, source ingestion, contact/privacy policy and streaming link delivery.
The technical architecture, simple call guide, tool guide, facts guide, privacy/submissions guide, editor operations and launch controls have been updated. Earlier unresolved agenda/profile/confirmation statements are superseded by the final evidence; historical test artifacts remain unchanged.
Deployment status: these AI fixes remain local and unpushed. The scoped published-page Pinecone refresh was applied; no Sanity documents or published launch/model settings were edited during this fix cycle. Final test apps were stopped, test namespaces were verified empty, and user dev servers were untouched. No real enquiries were sent.
The separate earlier event-route isolation repair was deployed to main at 6ba30a04: saved production smoke shows the unpublished Miami agenda paths returning 404 while Paris routes and event roots return 200. Its 11 isolated regressions are separate from the AI counts. Deployment evidence · Saved production smoke. No deployment was made during this AI fix/documentation cycle.
Report rebuild and artifact contracts · Final local verification · Cleanup. Rebuilding this report only reads retained evidence; it does not restart paid tests.
Evidence
All 185 reviewed occurrences · JSON review ledger · Latest retained 558 raw results · Scores/exclusion denominator · Per-round execution/latency measurements · Frozen final source manifest · Source refresh · Published/raw ingestion evidence · Miami publication check.