AI fixes — failure-only retest report
Index
Scope and status
Completed 37 unique previously failed scenario occurrences: 19 OpenAI GPT-5.4-mini and 18 Claude Sonnet 4.6. After reviewing the first results, 14 still-failing occurrences affected by additional fixes were rerun. A further nine unresolved occurrences were tested in round three, followed by three unresolved/error occurrences in round four. These are 63 executed scenario occurrences, not 63 distinct cases. Passing, partial and blocked baseline cases were excluded, as were cases reviewed as solely obsolete expectations. No new full-suite run occurred.
The user confirmed Gateway credit with the $40.21 screenshot. An initial suite-file wrapper error prevented acceptance execution; it was repaired. Accidentally copied old baseline files are archived as setup-error evidence and excluded from these totals. The completed Claude sponsorship failure scenario was retained and not repeated.
Both variants use the same profile extractor (OpenAI GPT-5.6-luna), narrow Jev selector, simulator and judge as the original benchmark. No production/staging website or AI configuration changed. At the user’s request, the existing shared Paris agenda and sitemap Pinecone namespaces were refreshed from their current published sources; see the source-refresh evidence below. All submission destinations are local capture receivers.
Original supplied UAT
This block covers only the original 95 cases, C01–C95 under each model. It excludes later additions, regression/workflow cases, scripted submission checks and FAQ freshness tests.
| Model |
Original raw failures |
Latest retained raw failures |
Latest pass |
Latest partial |
Latest blocked |
| openai |
19 / 95 (20.0%) |
12 / 95 (12.6%) |
47 |
33 |
3 |
| claude |
17 / 95 (17.9%) |
13 / 95 (13.7%) |
50 |
27 |
5 |
Latest retained combines targeted reruns with untouched original results. Seven original-UAT cases were rerun on OpenAI and five on Claude; some received subsequent failure-only corrections. We did not repeat the whole 95-case block.
Percentages use all 95 supplied cases as the denominator. Partial and blocked outcomes are separate from passes. Raw failures include obsolete expectations and judge evidence gaps, and are not a confirmed application-defect rate.
The 37 selected scenario occurrences below span original UAT and later blocks; do not add their totals to these 95-case totals. Count evidence · Original comparison.
Adjusted scores
These scores remove 33 model/block/case occurrences with confirmed obsolete expectations or unsupported evaluator judgments: 13 OpenAI and 20 Claude. They exclude those occurrences from the denominator; they do not relabel them as passes. The raw results and transcripts above remain intact. This is a conservative, itemised adjustment rather than a complete re-audit or a new test run.
Full passes only; partial and blocked outcomes remain in the denominator. Missing content, mixed failures and unresolved application defects are retained. The same exclusions apply to the before/after figures, so the comparison uses a consistent cohort.
| Scope |
Model |
Raw latest pass rate |
Excluded |
Adjusted before |
Adjusted latest pass rate |
| Original supplied UAT |
openai |
47/95 (49.5%) |
1 |
44/94 (46.8%) |
47/94 (50.0%) |
| Original supplied UAT |
claude |
50/95 (52.6%) |
5 |
48/90 (53.3%) |
50/90 (55.6%) |
| Original supplied UAT |
Combined |
— |
6 |
— |
97/184 (52.7%) |
| Overall benchmark |
openai |
163/279 (58.4%) |
13 |
151/266 (56.8%) |
163/266 (61.3%) |
| Overall benchmark |
claude |
177/279 (63.4%) |
20 |
168/259 (64.9%) |
177/259 (68.3%) |
| Overall benchmark |
Combined |
— |
33 |
— |
340/525 (64.8%) |
The full benchmark originally contains 279 occurrences per model; the original supplied UAT is a separate 95-case block per model. Adjusted totals must not be added to the raw totals. Variant denominators differ because some evaluator errors are model-specific; these percentages are descriptive, not a controlled model-ranking claim.
What stays in the score
- T02 on both models: historical pricing is wrong in the test, but the response also misses the required concrete pass-preference question.
- OpenAI K17: its old complaints-email expectation is wrong, but the unrelated duplicate-message warning is a real issue.
- Claude C97: final consent is now required, but the transcript has independent agenda/dispatch problems.
- C14 and C01 partials: some summary/qualification rules are stale, but independent plan completeness or eligibility issues remain.
- F08: the old script omits required company/role, but repetitive recovery and malformed-email handling are still concerns.
- C116: missing volunteer content is an unresolved content gap, not a pass.
- C18/C21, C33, OpenAI C54/C68/C83 and other unsupported-fabrication claims: evidence is not sufficient to exclude the whole current outcome without resolving event-scoping, source-publication or independent wording/behaviour questions.
- All other partial/blocked/unreviewed results remain. The separate one-day/two-day agenda defect is still recorded, even though H01 passes its handoff criterion.
Exclusion register
| Model |
Block / case |
Raw latest |
Exclusion reason |
Evidence |
| openai |
feedback / F07 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| claude |
feedback / F07 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| openai |
feedback / F09 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| claude |
feedback / F09 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| openai |
ticket-precision / F07 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| claude |
ticket-precision / F07 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| openai |
ticket-precision / F09 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| claude |
ticket-precision / F09 |
fail |
outdated prices: Only the old fixed 2,995/4,995 prices cause failure; the recorded structured badge source returns 3,195/5,495. Other evaluated flow and cancellation criteria were satisfied. |
Transcript · Source |
| openai |
ticket-precision / T01 |
fail |
outdated prices: Diamond answer matches the recorded structured price, rather than the fixed historical 4,995 expected by this test. |
Transcript · Source |
| claude |
ticket-precision / T01 |
fail |
outdated prices: Diamond answer matches the recorded structured price, rather than the fixed historical 4,995 expected by this test. |
Transcript · Source |
| openai |
uat-additions / C98 |
fail |
conflicting wording: The blanket no-British-spelling criterion forbids the current approved app-owned personalised CTA. Rewrite the style expectation before scoring this case. |
Transcript · Source |
| claude |
uat-additions / C98 |
fail |
conflicting wording: The blanket no-British-spelling criterion forbids the current approved app-owned personalised CTA. Rewrite the style expectation before scoring this case. |
Transcript · Source |
| openai |
uat-additions / C104 |
fail |
outdated roster: The frozen August six-speaker note contradicts the seven keynote-format sessions in the current live source. Latest answers match that source. |
Transcript · Source |
| claude |
uat-additions / C104 |
fail |
outdated roster: The frozen August six-speaker note contradicts the seven keynote-format sessions in the current live source. Latest answers match that source. |
Transcript · Source |
| openai |
knowledge / K09 |
partial |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| claude |
knowledge / K09 |
partial |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| openai |
knowledge / K18 |
fail |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| claude |
knowledge / K18 |
fail |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| openai |
knowledge / K21 |
partial |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| claude |
knowledge / K21 |
partial |
outdated contact policy: Grading penalises support@unleash.ai instead of the former Customer Success/events contact route. Current policy uses Support for these non-SPEX requests. |
Transcript · Source |
| openai |
knowledge / K20 |
fail |
outdated privacy policy: Old criteria require a privacy-page redirect and deny access; the current explicit policy says the bot cannot share stored information and directs the visitor to support@unleash.ai. |
Transcript · Source |
| claude |
knowledge / K20 |
fail |
outdated privacy policy: Old criteria require a privacy-page redirect and deny access; the current explicit policy says the bot cannot share stored information and directs the visitor to support@unleash.ai. |
Transcript · Source |
| claude |
knowledge / K17 |
fail |
outdated contact policy: The failure is solely the superseded complaints email expectation. OpenAI is retained because its sign-off also hits the duplicate-message warning. |
Transcript · Source |
| openai |
uat-additions / C97 |
fail |
outdated consent policy: The raw failure penalises obtaining explicit final consent after the name arrives; that confirmation is required. Claude is retained for its separate missing agenda/dispatch problems. |
Transcript · Source |
| openai |
original-uat / C51 |
fail |
judge source-evidence gap: The judge cannot verify prices/benefits from its context; the recorded structured badge source contains the quoted prices and tier summaries. Missing judge evidence is not demonstrated invention. |
Transcript · Source |
| claude |
original-uat / C51 |
partial |
judge source-evidence gap: The judge cannot verify prices/benefits from its context; the recorded structured badge source contains the quoted prices and tier summaries. Missing judge evidence is not demonstrated invention. |
Transcript · Source |
| claude |
original-uat / C78 |
fail |
judge source-evidence gap: Recorded badgesLookup and FAQ observations substantiate the prices and conditional onsite policy penalised as invented. |
Transcript · Source |
| claude |
original-uat / C54 |
fail |
judge source-evidence gap: The fabrication allegation is contradicted by the case-specific retrieved package, startup story or transfer-policy evidence. This does not certify ongoing source freshness. |
Transcript · Source |
| claude |
original-uat / C77 |
fail |
judge source-evidence gap: The fabrication allegation is contradicted by the case-specific retrieved package, startup story or transfer-policy evidence. This does not certify ongoing source freshness. |
Transcript · Source |
| claude |
original-uat / C80 |
fail |
judge source-evidence gap: The fabrication allegation is contradicted by the case-specific retrieved package, startup story or transfer-policy evidence. This does not certify ongoing source freshness. |
Transcript · Source |
| claude |
knowledge / K04 |
fail |
judge source-evidence gap: The retrieved Terms & Conditions contain the refund tiers and service charge; current Support routing supersedes the old contact criterion. OpenAI remains for its independent incomplete terms explanation. |
Transcript · Source |
| claude |
workflow-router / W06 |
partial |
judge summary-evidence gap: The visible seeded agenda, current recipient summary, explicit confirmation and single accepted agenda-email capture satisfy the current summary requirement; reprinting the entire agenda was not required. |
Transcript · Source |
| claude |
flow-selector / W06 |
partial |
judge summary-evidence gap: The visible seeded agenda, current recipient summary, explicit confirmation and single accepted agenda-email capture satisfy the current summary requirement; reprinting the entire agenda was not required. |
Transcript · Source |
Machine-readable scores and exclusions · OpenAI complete inclusion register · Claude complete inclusion register.
The adjustment does not change application code, expected-case files, deployment, live sources or historical verdicts. Fixed historic prices, roster notes and policy assertions should be rewritten from current source/policy contracts before their next execution; no fresh rerun occurred for this scoring change.
Missing content percentages
Confirmed minimum, based on reviewed source evidence: one missing content area—current Paris volunteering eligibility/application information and the five-team role breakdown—affects two retained case occurrences, C116 on OpenAI and Claude. The published-page snapshot and uploaded FAQs do not provide the required answer. This is a content/publication gap; the current replies safely provide Support rather than inventing roles or linking a removed page. Raw failures stay in the scores.
| Scope |
Model |
Confirmed missing-content cases / scored cases |
Percentage of scored cases |
Share of retained raw failures |
| Original supplied UAT |
OpenAI |
0 / 94 |
0.0% confirmed |
0 / 11 (0.0%) confirmed |
| Original supplied UAT |
Claude |
0 / 90 |
0.0% confirmed |
0 / 9 (0.0%) confirmed |
| Overall adjusted benchmark |
OpenAI |
1 / 266 |
0.4% |
1 / 41 (2.4%) |
| Overall adjusted benchmark |
Claude |
1 / 259 |
0.4% |
1 / 35 (2.9%) |
| Overall adjusted benchmark |
Combined |
2 / 525 |
0.4% |
2 / 76 (2.6%) |
| Failure-only retest, after exclusions |
Combined |
2 / 31 |
6.5% |
2 / 2 (100.0%) |
The original supplied UAT is C01–C95; C116 belongs to the later additions. “0 confirmed” does not mean no missing content exists in the original UAT: the remaining outcomes have not all received a complete source audit. Overall raw denominators give 1/279 (0.4%) per model and 2/558 (0.4%) combined. The retest row covers only the 37 deliberately selected failures, of which six are excluded as outdated/evaluator-only; its 100% figure must not be applied to the full benchmark's remaining failures.
Not counted as confirmed content-only gaps: C18/C21 event-scoping/source concerns, C33 volunteer publication concerns, C24/C26 requested articles/webinars/reports being substituted, and C85 unsupported attendee-industry claims. These may involve source freshness, retrieval or response behavior, and require case-specific source review before assigning a content-only cause. A missing answer in a transcript is insufficient evidence that the content itself is missing. Older volunteer outcomes are not automatically reclassified from the later source snapshot.
Action: editorial owners should confirm whether a current volunteer program should be published; if so, publish approved eligibility, application link and team/role details, sync those pages/FAQs into the existing Pinecone namespaces, and rerun C116 on both models. If the program is intentionally unavailable, update the expected behavior to the approved Support fallback. Do not fabricate or restore unpublished content merely to satisfy a historical test.
Evidence: OpenAI C116, Claude C116, published source snapshot, source-refresh results, adjusted denominators. These percentages describe test occurrences, not the percentage of site pages or knowledge-base documents missing. No new tests or paid calls were run for this breakdown.
Results
All selected baseline verdicts were fail. Raw judge labels below are preserved; they are not a reviewed application-defect count.
| Variant |
Selected cases |
Round 1 pass / partial / fail |
Latest pass / partial / fail / blocked / error |
| openai |
19 |
4 / 4 / 11 |
12 / 4 / 3 / 0 / 0 |
| claude |
18 |
7 / 2 / 9 |
9 / 6 / 3 / 0 / 0 |
OpenAI case detail · Claude case detail · Raw transitions CSV · Detailed transcript review
Verified fixes and remaining findings
The latest retained raw totals are 21 pass / 10 partial / 6 fail, versus 37 selected baseline failures. This is a targeted before/after result, not a full-system pass rate or causal model comparison.
Verified fixes include natural ticket confirmation, saved agenda email with one correct handoff, ticket navigation, failed-sponsorship acknowledgement, headcount handling, retries, stored email masking, complete free-ticket FAQ answers, stand-plan/manual contact routing and conditional onsite registration. The Claude registration detour now collects the selected pass and saved contact details, shows a single consent summary and captures one delegate enquiry after confirmation.
The six remaining raw failures are reviewed separately: two keynote cases use obsolete August ground truth; two volunteer cases require a page currently absent from the published sources; OpenAI C98 uses a stale spelling expectation; Claude C78’s evaluator lacked actual pricing/FAQ evidence. Raw verdicts are preserved. Partial outcomes and broader UAT failures still require separate review; no complete-system certification is claimed.
Source refresh and final focused fixes
The user requested refreshing the existing source-to-Pinecone pipeline instead of moving runtime lookups to event-management/Sanity APIs. Runtime agenda lookups remain in Pinecone. The new exact read-only eventFacts lookup also reads FAQ/uploads and complete event-page bodies from existing Pinecone namespaces.
The refresh wrote 237 Paris agenda vectors and 145 page vectors covering 33 currently published Paris pages, removing 34 vectors belonging to pages no longer present. Before/current-source/after artifacts are retained in refresh evidence. Other event and knowledge namespaces, website deployment and AI switches were unchanged.
The refreshed live event-management source itself lists seven keynote-format sessions, including Brian Glaser and Stephen Childs, while Eric Mosley is absent from its keynote set. The older local agenda JSON and August six-speaker judge note are stale relative to that source. The earlier suspicion that Pinecone alone had the wrong formats was not supported by the live source check. Both C104 answers matched the current source; their raw failures remain because the original judge note was deliberately preserved. Current source snapshot.
No current volunteer page or volunteer FAQ answer is available. The two C116 replies correctly avoid inventing five team names or linking the removed page, and provide Support instead. Restoring that public content is an editorial decision, not something these tests should force into the answer.
OpenAI R01/R02 passed with complete HR programme, group-ticket and sponsor-package guidance. R07 initially produced an empty response observed by the runner, with the underlying cause not established; a focused retry passed with manual instructions and Customer Success. OpenAI C78 initially received a blocked judge label because the evaluator lacked policy evidence; the final retest included the captured current FAQ in evaluator notes, unchanged criteria, and a published Attend link, and passed. These evidence changes are recorded in round-four manifest.
Claude H01 initially reached the ticket workflow but repeatedly requested quantity because model extraction could erase saved booking fields. The fix preserves existing booking values through null extraction, recognises an explicit personal “a ticket” request as one, and clears old fields on event/new-booking changes. The final fixed six-message replay passed, with a contact summary, explicit consent and one HTTP 200 local capture. It does not prove real inbox or Salesforce delivery. The transcript’s initial agenda still spans two days despite the visitor’s one-day request; that separate plan-duration issue is recorded rather than hidden by the handoff verdict.
Final H01 transcript · Onsite transcript · Stand plans transcript · Implementation guide.
Final fix outline and next actions
Implementation and source changes
| Area |
What changed |
Where it reads or acts |
Verification / limit |
| Precise FAQ answers |
Added the read-only eventFacts tool for free tickets, onsite registration, stand plans/manual access and volunteering; narrow questions can return complete FAQ answers directly instead of partial search snippets. |
Existing general/event FAQ-upload namespaces and event-page bodies in the Pinecone sitemap namespace. |
Event isolation, FAQ aliases, conditional wording, contact routing and missing-source controls tested. No new external plugin or runtime skill installed. |
| Internal links |
Exact fact answers filter internal destinations against the synced event-page catalogue. |
Pinecone page metadata refreshed from published Sanity sources. |
Removed/unpublished destinations are omitted; missing information falls back to Support. Correctness depends on the catalogue being kept current. |
| Keynotes |
The explicit session Format determines keynote status, case-insensitively; contradictory older markers and speaker prominence do not override it. Explicit roster requests read the complete agenda. |
Existing Pinecone agenda namespace. |
Shared lookup tests passed. Current source has seven keynotes, so the old six-name test note must be updated independently of code. |
| Agenda → tickets |
Explicit registration help or proceeding with a named pass switches to the existing ticket workflow. An email in that workflow no longer routes into agenda delivery because an earlier plan exists. |
Existing workflow state and transcript guards. |
Claude H01 fixed-message replay passed; the selector still chooses destinations and does not grant consent. |
| Saved booking details |
Null extraction no longer erases saved quantity/pass/event details. An explicit personal request for “a ticket” supplies quantity one; event/new-booking changes reset old details. |
Existing profile-extraction and ticket-enquiry flow. |
Unit controls include corrections, event changes, new bookings, negation and factual badge questions. H01 shows one ticket and the saved contact details before consent. |
| Confirmation and delivery |
Retained the contact summary, privacy/terms notice, explicit confirmation, cancellation and delivery-result checks. “Looking forward to the event” no longer gets mistaken for a malformed agenda recipient after a send confirmation. |
Existing ticket/SPEX/agenda handoff infrastructure. |
One H01 HTTP 200 local capture after confirmation; no real email or CRM submission. |
| Tool and production controls |
The new read-only tool uses the existing capability registry and can be disabled through the Sanity configuration. |
Existing tool filters and production/email master controls. |
Capability tests passed. No production/staging AI setting or CMS deployment changed. |
| Content refresh |
Re-uploaded Paris agenda and currently published Paris pages using existing sync functions; removed stale page vectors. |
Shared existing Pinecone namespaces, not a new answer source. |
237 agenda vectors, 145 page vectors / 33 pages, 34 removed vectors. Before snapshot, current published sources, refresh result. |
Deployment status: the application fixes are local and have not been pushed or deployed. The requested content refresh has already been applied to the existing shared Pinecone data. Source-fetching remains in the publishing/refresh pipeline; runtime fact and agenda lookups remain in Pinecone.
The agenda has an existing ten-minute cache per warm application instance. Upload clears the cache in the process performing the upload; that does not establish immediate invalidation of every other instance. The isolated test apps were restarted before the final retests. The 38.22-second refresh duration measures the bulk refresh operation, not FAQ-to-answer propagation or an SLA.
Focused case results
| Case / variant |
Latest raw verdict |
Reviewed behaviour and evidence |
| R01 / OpenAI |
Pass |
Complete HR professionals programme, group-ticket and sponsor-package guidance. Transcript. |
| R02 / OpenAI |
Pass |
The alternate free-ticket phrasing reaches the same complete FAQ guidance. Transcript. |
| R07 / OpenAI |
Pass |
Stand plans route to the exhibitor manual and customersuccess@unleash.ai, without starting sponsorship sales. An earlier attempt produced an empty reply observed by the runner; its underlying cause was not established. The focused retry passed. Transcript. |
| C78 / OpenAI |
Pass |
Onsite availability stays conditional on capacity and links to Attend. Current FAQ evidence was added to evaluator notes, with the acceptance criteria unchanged. Transcript. |
| H01 / Claude |
Pass |
Agenda-to-registration detour reaches a contact/quantity/pass summary, explicit consent and one accepted local delegate enquiry. Transcript. |
| C104 / OpenAI and Claude |
Fail / fail |
Answers match the seven current keynote-format sessions. The original August six-speaker judge note is obsolete. Raw verdicts remain intact. OpenAI, Claude. |
| C116 / OpenAI and Claude |
Fail / fail |
The expected volunteer page/team list is currently absent. Replies provide Support and avoid invented roles or a removed link. OpenAI, Claude. |
OpenAI C98 and Claude C78 were not repeated in these final rounds: the former has an obsolete spelling expectation; the latter's prior failure conflicts with the recorded badge/FAQ tool evidence. Together with C104/C116, these account for the six remaining raw failures in the 37 selected occurrences. This does not resolve every partial outcome or every untouched original-UAT failure.
Outstanding work and operational limits
| Item |
Category |
Next action |
| H01's initial plan spans two days despite a one-day request |
Observed application defect, separate from its passing ticket-handoff criterion |
Enforce the requested duration/day in agenda construction and add a focused regression. Not fixed in this batch. |
| Five-team volunteer answer cannot be sourced |
Content/publication gap |
Editorial team must decide whether to restore the volunteer page and approved team/role information, then sync and retest. Do not recreate unpublished content merely to satisfy the old test. |
| Six-keynote August ground truth |
Obsolete expectation |
Update the benchmark source note from the current agenda, preserving the historical run and raw labels. |
| OpenAI C98 spelling expectation |
Obsolete expectation |
Align the expected wording with the approved current CTA before a future retest. |
| Claude C78 evaluator/source mismatch |
Evaluator evidence gap |
Supply actual structured badge/FAQ evidence to future evaluation; retain its historical raw failure. |
| Ten partial outcomes and untouched original UAT failures |
Separate review scope |
Review individually before claiming wider release readiness. Only the selected failures were rerun here. |
| FAQ edits reaching fresh and existing conversations |
Untested operational check |
Run the six prepared isolated freshness tests and measure propagation; no timing claim has been established. |
| Real inbox delivery, Salesforce processing, deployment and production controls |
Untested operational checks in this batch |
Validate independently when deployment is authorised. Local HTTP 200 capture proves only the test receiver accepted the request. |
| Repository typecheck errors outside the changed files |
Existing verification limitation |
Resolve separately; do not describe this as a clean whole-repository typecheck. |
Implementation guide · Round-three frozen code · Round-four frozen code and evaluator-evidence changes.
Changes
| Confirmed defect |
Implemented correction |
Relevant cases |
| Ordinary affirmative send replies fail consent matching |
Accept constrained natural confirmation wording, including “Yes, please send it through” and “Yes, that's all correct. Please send it”; reject conditions, postponement, cancellation and unrelated follow-ups |
C01, C03, C14, C17 |
| Email delivery triggers a privacy refusal |
Distinguish delivery to the visitor's email from requests to disclose/export stored personal information |
C66, C98 |
| Employee counts are treated as historic event years |
Exclude year-shaped headcount/quantity/currency values followed by explicit units from the past-event guard |
F04, F13 |
| Exhibitor stand plans use general Support |
Recognize stand-plan/build/design/contractor and exhibitor-manual questions as clear exhibitor enquiries; retain privacy and ticket precedence |
R07 |
| Saved agenda email is requested again |
Recognize “to my email” as a reference to an existing saved address, preserving malformed explicit-address validation |
W06 |
| Old agenda mode interferes with a ticket/SPEX switch |
Clear agenda-email pending status for ticket/SPEX destinations, remove agenda tools outside the agenda destination, and select the handoff tool from the current workflow state |
H01, W03, W10 |
| Consent makes submission tools available on unrelated turns |
Expose submission tools only for the current confirmed handoff; an acknowledgement or factual side question cannot invoke them simply because session consent is stored |
sponsorship-failure, FLOW-SIDE |
| Thanks after a failed send causes an automatic retry |
Answer short acknowledgements after the authoritative failure directly, without model tools or another handoff |
sponsorship-failure |
| Corrected recipients are absent from the final summary |
Require the current name, work email, company and role in the confirmation summary; otherwise issue a fresh application-owned summary |
W07, W10 |
| Browsing lists with overlaps are promised for email and then rejected |
Convert an agenda email request without a validated displayed plan into a source-built, non-overlapping personal plan and contact summary before submission |
C107, F03, F14 |
| Logistics/upgrade questions prematurely qualify a buyer |
Exclude venue/day-of-event purchasing, purchase navigation and existing-registration upgrades from deterministic new-booking intent |
C78, C81, C108 |
| Requested content is substituted or unsupported taxonomy inferred |
Strengthen current-turn instructions for articles/webinars/reports, attendee industries and volunteer taxonomies; require explicit source evidence for hosts/keynote formats and prohibit placeholder URLs |
C24, C26, C85, C95, C104, C116, R14 |
| SDK 7 tool/generation telemetry bypasses PII masking |
Mask all exported span attributes before Langfuse export, including gen_ai.* messages and tool arguments, retaining public UNLEASH contacts and numeric usage |
Synthetic telemetry evidence |
| Empty response retries hit the duplicate-message warning |
Count delivered exchanges within the current chat instead of failed HTTP attempts; keep the shared IP rate budget and submission safeguards |
R21, R26 |
| Five ordinary form tests instantiate native inputs illegally |
Create fixture inputs through document.createElement('input') under the configured happy-dom environment |
formBuilderConditions tests |
Additional corrections preserve full natural agenda-email acceptance through later contact collection and produce an application-owned contact summary before a send. Factual ticket navigation/venue/upgrade, free-ticket eligibility and exhibitor logistics questions are protected from new-lead qualification even when the selector chooses a commercial destination.
Local verification
- Latest owning-package checks: 231 web tests passed / 0 failed, and 36 shared AI tests passed / 0 failed. Form tests previously passed 94/94; the full ordinary monorepo suite was not repeated.
- After the second corrections, the focused flow test file passed 19/19, including natural email acceptance, negative/delayed controls and commercial-information versus actual buying requests.
- Biome passed on all 14 final touched source/test files. Web typecheck still fails elsewhere; no diagnostics matched the final changed files. Typecheck log.
- Original ordinary results remain in the baseline ordinary report, with a follow-up link.
Latency, streaming and handoffs
Metrics use the latest retained result for each selected case. Visitor first-token time includes profile extraction. They describe small, selected local concurrent batches and are not an overall model benchmark.
| Variant |
Measured turns |
First token p50 / p90 |
Profile p50 |
Multi-chunk responses |
Accepted local / attempts |
| openai |
59 |
3.67s / 6.96s |
2.16s |
19 |
5 / 5 |
| claude |
67 |
3.71s / 9.26s |
2.17s |
16 |
9 / 10 |
Multi-chunk answers confirm streaming continues in these cases; deterministic short replies use one chunk. HTTP 503 is rejected delivery, never successful submission. Captures remain in each transcript. Compiled metrics also separate all attempts from latest selected results.
The synthetic post-fix Langfuse submission trace masks the email in persisted tool arguments. Names, company and role remain visible; this is not full anonymisation. The existing broad phone masking also masks parts of session timestamps. Stored trace evidence.
FAQ freshness
Six new isolated operational tests are prepared for FAQ create/edit, existing-chat refresh, shorter-upload replacement, deletion and empty-FAQ cleanup while retaining uploads. They were not run under this failure-only instruction; no propagation time or SLA has been measured. Test method. The empty FAQ sync cleanup defect is corrected locally, with the live verification still pending.
Evidence and limits
- Raw original benchmark, all four retest rounds remain separate. Baseline and later runs use adaptive follow-ups, not identical conversation replays.
- The source snapshots are recorded in round-one manifest and round-two manifest, round-three manifest, and round-four manifest.
- Transcript counts and selected case identifiers were checked; no stale setup-error files contribute to totals.
- Real inbox delivery, Salesforce processing, shared Sanity publishing, live FAQ propagation, production kill-switch changes and all previously passing cases remain outside this retest.
- No statistically significant model improvement, full release certification or actual spend total is claimed. The credit screenshot was a point-in-time balance.
- No code push, deployment, real email or CRM write was made.