AI fixes and verification — 5 October 2026

Index

Status and scope

The failure-only review is complete. No verified application defect remains open in the reviewed 185 model/block/case occurrences. This does not mean every raw judge verdict passed, or that all possible conversations have been tested.

The AI fixes reviewed here are local, not pushed or deployed. The final tested source is base b2534f60 plus the SHA256-listed files in the final manifest. Existing production settings were not changed. Five published Paris pages were refreshed through the existing Pinecone pipeline; no Sanity documents were edited. Test submissions went only to loopback receivers. Isolated test apps were stopped; the user's development servers were left alone.

The compiled report preserves raw results alongside transcript/source review. This document explains the implementation changes, their evidence and their limits. September comparisons and the original October benchmark remain historical evidence; they are not silently replaced by targeted reruns.

How the system now handles a message

  1. The widget supplies the message, page URL, recent conversation and saved profile. The page/conversation determines Paris or Miami; an ambiguous event is clarified in ordinary conversation. There are no event/workflow selectors.
  2. Profile extraction merges explicit new visitor facts. Missing extraction values do not erase known company, role, name or email. A new event or booking can reset booking-specific details.
  3. The API checks visibility/capabilities, cancellation, request boundaries and narrow factual policies. Read-only information requests do not start contact qualification.
  4. An active workflow continues directly for routine replies and recognized send confirmations. Otherwise Jev selects the existing tickets, SPEX, agenda or information destination, with fallback on uncertainty/error.
  5. The conversational model receives the known-facts table and permitted tools. Structured lookups and deterministic helpers handle exact facts, ticket guidance and agenda planning where appropriate.
  6. A submission requires a current request summary, valid contact information, affirmative confirmation and fresh capability checks. The tool outcome determines the delivery message.
  7. Ordinary model text streams. Exact factual replies, validated plans and handoff outcomes may arrive as complete server-owned text. Langfuse review happens after delivery.

Jev still selects a destination. It does not generate or validate answers, authorize consent, or send requests. This cycle improved workflow execution and source grounding rather than adding a second answer-validation model.

Fix register

Each row describes an observed problem addressed in this cycle. Regression files test the relevant contracts; the report links individual model transcripts and source evidence.

Area / observed problem Resulting behavior Code / regression evidence
Information accidentally starts sales qualification Pricing, logistics and polite requests to hear about pass options stay informational. Explicit help buying/contact requests can start tickets. aiFlows.ts, tests
Routine replies get reclassified or repeat an offer Active-flow replies and recognized confirmation skip Jev; explicit changes and factual detours preserve the correct context. flowSelector.server.ts, workflow guards
Cancellation or failed delivery resumes collection Withdrawal stops the pending flow. A rejected send ends the handoff and offers Support; acknowledgements do not reopen it. A new attempt requires an explicit request. workflows.ts, API, tests
Natural confirmation loops Bounded affirmative wording is accepted after a handoff question, including a group quote confirmation. Negative, conditional and unrelated field-change text is not send authorization. workflows.ts, passing quote replay
Saved details or consent cause premature submission SPEX and tickets show the current contact/request summary and wait for confirmation even with saved consent. The visitor's commercial objective is preserved. salesConfirmation.ts, tests
Invalid/corrected email, repeated profile questions Invalid format is addressed before further qualification; escaped @ is normalized; corrections and known fields survive extraction. No name/company is inferred from an email domain. handoffContact.ts, userProfile.ts, profile route
Estimated group size becomes an invented exact count Explicit quantity ranges remain estimates. Explicit later counts replace them; new bookings do not inherit stale pass/count choices. userProfile.ts, tests
Group quote requires choosing a pass first A requested quote retains “Pass options / group quote (no pass selected)” and can proceed without inventing a pass or Miami price. ticketAnswer.ts, quote replay
Pass-help loop / stale prices Repeated help advances to a recommendation/choice. Numeric prices come from current structured pricing; Explorer remains “Price on application.” ticketAnswer.ts, livePricing.server.ts, tests
One-day agenda expands into two days / overlaps Explicit day and source year, eligibility, non-overlap and transfer time are enforced in the planner. agendaPlan.ts, tests
Agenda email loses interests, recipient or selection Saved fields survive; corrected recipients are reused; browsing selections are validated into a plan before handoff. Email content/URLs are server-owned. chatDelivery.ts, agendaEmail.ts, tests
Stage/speaker/keynote answers use the wrong scope Exact stage matching avoids Stage 1/10 confusion. Current session Format identifies keynotes. Missing named speakers are described as unconfirmed, not categorically absent. agendaListing.ts, agenda.ts, tests
FAQ conditions disappear / fragments substitute for policy eventFacts loads complete matching Q/A entries and existing indexed page bodies. It preserves conditions and withholds unpublished facts. eventFacts.ts, loader, tests
Startup information repeats or uses the wrong application Eligibility, package inclusions, networking-badge checkout and the separate Award application have distinct responses. Published audience is not eligibility; applying is not approval. eventFacts.ts, exact replay
Unsupported discount/recording entitlement Missing promotional prices stay unknown. Terms about separately purchased recordings do not establish inclusion with a badge. eventFacts.ts, source evidence
Withdrawn FAQ/file facts are invented Recognized narrow named facts require current evidence for the same entity and requested field; related entities and old history cannot supply a replacement. specificKnowledgeFact.ts, tests, four live probes
Sponsor names missing from indexed pages Ingestion includes public referenced names and explicit image alt text; related reference publishes refresh affected pages. Asset filenames are not approved sponsor names. pageCatalogue.ts, sync.ts, published/raw verification
Media questions return the wrong format Report, research and webinar requests retain their requested format. Latest articles continue to use date-aware Sanity retrieval. mediaSearch.ts, tests
Wrong contact route / ordinary details mistaken for export General/ticket/refund/privacy/failure uses Support. Clear sponsor/exhibitor/programme questions use Customer Success. Volunteering contact details is not a stored-data export request. aiResponsePolicy.ts, tests
Bad links, missing external hotel URL, joined paragraphs Source booking destinations are surfaced; quoted links are normalized, unsafe/empty destinations become plain text, and model step boundaries remain readable. hotelBooking.ts, markdownDelivery.ts, chatDelivery.ts

Routing, memory and confirmation

The prompt receives a known/missing values table. Browser session storage and process-local workflow state remain separate; neither is authenticated identity or a durable shared consent database. Company/role extraction does not derive values from an email address. Quantity ranges use a dedicated field instead of silently rounding.

Actual current extract from userProfile.ts:

const pass = requestsTicketQuote(message)
    ? 'Pass options / group quote (no pass selected)'
    : proposed.ticketPass || (changed ? undefined : current.ticketPass)

The active-flow fast path includes affirmative confirmation so a valid send reply does not become a new classification problem. Actual extract from flowSelector.server.ts:

if (
    input.active &&
    (routineContinuation(input.message) ||
        confirmedHandoff(input.history, input.message))
)
    return selectDestination('continue', 1, input.active)

Confirmation is scoped to the preceding handoff question. The recognizer first rejects negative/conditional wording; accepted natural language remains bounded rather than treating any sentence starting with “yes” as permission. Actual extract from workflows.ts:

if (
    /\b(?:not|no|never|cancel|withdraw|revoke|if|unless|later|tomorrow)\b|don['’]t/i.test(
        message,
    )
)
    return false
if (!last || !isHandoffQuestion(last.text)) return false

Saved session consent survives a valid email correction. It does not bypass review of a new enquiry. Cancellation creates a history boundary so an earlier affirmative cannot authorize a later request. A rejected submission receives an honest failure message and ends qualification; there is no automatic retry or claim of successful delivery.

Source facts and knowledge updates

The retrieval architecture is preserved. FAQ and event-page fixes use the existing general/current-event/sitemap Pinecone stores. Editors publish or update approved source content, then refresh those stores; the assistant does not switch to an unrelated runtime source to compensate for stale indexing.

eventFacts is a read-only tool in the shared capability registry. It now covers 23 topic identifiers: free-ticket pathways, onsite registration, stand plans, volunteering, refunds, transfers, upgrades, invoice/payment, floorplans, registration, startup eligibility/options/inclusions/Award, attendance, industries, recordings, discounts, HR Leader eligibility, Explorer applications, parking, opening hours and sponsorship options. Recognition is narrow and source-dependent; this is not exhaustive coverage of every possible policy question.

The loader reconstructs overlapping chunks into complete documents and selects complete Q/A aliases scoped to the current event. Exact replies preserve conditional availability. Indexed page links are filtered against the published catalogue. Startup Award wording comes from its published competition page; the support-package value is not represented as a cash prize, guaranteed selection or badge discount.

The withdrawn-fact check is deliberately conservative and narrow. It recognizes certain proper-name room/document questions and requires evidence for both entity and field. Actual extract from specificKnowledgeFact.ts:

return blocks.some((block) => {
    // FAQ question and answer must remain together; a different alias cannot lend its field.
    const normalized = normalize(block)
    const words = new Set(normalized.split(' '))
    return (
        normalized.includes(normalize(fact.entity)) &&
        fact.attributes.every((word) => words.has(word))
    )
})

It is not a universal semantic fact checker. Ordinary retrieval answers still depend on model grounding and current sources; tests establish the specific removed-room and removed-file behaviors, not zero hallucination across all topics.

Other authorities remain appropriate to their jobs: current numeric prices use published mappings and Inwink products; badge inclusions use published content; latest articles/upcoming webinars use date-aware Sanity queries; sessions use the complete stored agenda. Price/product caches are 60 seconds and agenda warm-cache duration is ten minutes, in addition to upstream/indexing freshness.

The October price sample verified General at €3,195 + VAT, Diamond at €5,495 + VAT, and Explorer as Price on application. These are dated evidence, not permanent constants. Historical €2,995/€4,995 assertions were obsolete; the code must keep using structured current prices.

The earlier focused refresh wrote 237 Paris agenda vectors and 145 page vectors for 33 published Paris pages, removing 34 vectors for pages no longer present. Earlier refresh evidence is retained separately. Five approved Paris page refreshes produced 22 vectors, with no stale-tail deletion needed in that refresh. Publication scope equivalence was checked on 41 documents in both published and raw perspectives, with 13 public referenced entities. This proves that ingestion change's observed scope; it does not prove every nightly or publish-triggered sync is operating correctly.

The planner enforces requested days/source year, usable source sessions, access restrictions, non-overlap, room-transfer time and bounded selection. Alternatives are replacements rather than extra simultaneous sessions. Browsing lists are validated before being converted into an email plan. Name/company/role/email corrections use the same profile retention and contact checks as other handoffs.

Session flyout links use agenda-schedule with the source event/session IDs. The earlier one-day-to-two-day defect and profile/email follow-up failures have been fixed and retested; previous documents describing them as unresolved are historical. Complete browser close/navigation behavior still needs a deployment smoke check.

Ordinary generated text continues streaming. Link normalization buffers only the current Markdown link, capped at 4 KiB; it does not hold the entire answer or ask Jev to approve it. Actual extract from markdownDelivery.ts:

const raw = match[2].trim().replace(/^["'<]+|["'>]+$/g, '')
const href = /^mailto:[^\s@]+@[^\s@]+\.[^\s@]+$/i.test(raw)
    ? raw
    : sanitizeHref(raw)
return href ? `[${match[1]}](${href})` : match[1]

An invalid/empty/placeholder URL becomes a readable label. A source URL is not invented. For hotels, the direct maintained booking destination is preferred over merely pointing to the internal travel page. Multi-step output preserves paragraph boundaries. Complete deterministic factual/agenda replies can legitimately be single-chunk responses.

Submissions, cookies and privacy

The existing production master switch, saved functional-cookie widget gate and independent submission/tool switches remain in place. The registry now contains 18 tools, including the read-only eventFacts. Disabled submission tools prevent starting/resuming qualification and are checked again before dispatch; read-only help can remain available. The pre-prod Preview override for demo submissions is separate from production permission. This cycle did not change published settings.

Current handoffs require known event, valid required contact fields, allowed intent, current summary and affirmative send confirmation. Agenda submissions additionally validate sessions and create Markdown/HTML from server-owned content. Enquiry consent is separate from cookie acceptance and marketing subscription. Generic leads carry opt_in: false; agenda opt-in is separate and defaults false.

Only an explicit captured 2xx response is webhook acceptance. HTTP 503 is a rejected attempt, even if the receiver recorded the payload. Webhook acceptance does not prove Salesforce persistence, Postmark delivery, SMTP acceptance or a paid ticket. Failure uses support@unleash.ai and does not automatically retry. There is no application guarantee of exactly-once downstream delivery.

Privacy/data-processing or stored-information requests are referred to support@unleash.ai. Ordinary provision/correction of contact details is allowed, including displaying those details in the agreed pre-send summary. This is not a data export/deletion service or proof of authenticated identity. Telemetry redaction is partial; names and free text can remain. Cookie withdrawal does not retroactively erase provider/CRM records or recall an accepted request.

See the privacy and submission guide for permission distinctions, payloads, tracking signals and retention boundaries.

Original supplied UAT

The supplied UAT is C01–C95 only: 95 cases per model, 190 model/case occurrences combined. Later additions, regression cases, focused replays and FAQ probes are separate.

Model Raw before today's fixes Latest retained raw Latest partial / fail / blocked
OpenAI 44/95, 46.3% 82/95, 86.3% 11 / 1 / 1
Claude 48/95, 50.5% 81/95, 85.3% 8 / 5 / 1
Combined 92/190, 48.4% 163/190, 85.8% 19 / 6 / 2

With the same itemized evidence-only exclusions applied before and after, the combined adjusted score is 92/171 (53.8%) → 163/171 (95.3%). Latest adjusted scores are OpenAI 82/85 (96.5%) and Claude 81/86 (94.2%). Content gaps, partial results and untested operational checks are not converted into passes.

These are rolling retained results: targeted reruns replace corresponding old results while untouched cases retain their original result. They are not a fresh execution of all 95 cases after the final change.

Entire retained benchmark

The complete original benchmark has 279 occurrences per model, 558 combined. It includes supplied UAT, later additions and regression blocks. The report keeps each block identifiable.

Model Raw before Latest retained raw Latest adjusted
OpenAI 151/279, 54.1% 239/279, 85.7% 239/244, 98.0%
Claude 168/279, 60.2% 239/279, 85.7% 239/247, 96.8%
Combined 319/558, 57.2% 478/558, 85.7% 478/491, 97.4%

There are 67 evidence-only exclusions from the adjusted denominator: 33 previously reviewed plus 34 newly reviewed. Combined adjusted before was 319/491 (65.0%). Every exclusion and denominator is recorded; raw verdicts are preserved. An obsolete expected price/address and an evaluator-only penalty are different from an application defect or missing content.

The comparison is descriptive. Adaptive selection, stochastic visitor/model/judge behavior and content refreshes prevent a statistical improvement or model-ranking claim from this run.

Failure-only executions and focused replays

The remaining cohort contained 185 previously non-passing occurrences. 177 received actual paid failure/partial-only retests; six confirmed volunteer-content occurrences and two browser/session operational checks were retained without retries.

Review outcome Occurrences
Raw pass / observed resolution 138
Obsolete expectation or evaluator-only issue 34
Content gap 10
Operationally unverified 3

Raw cohort grades remain 138 pass, 27 partial, 18 fail, 2 blocked. The reviewed categories explain those grades; they do not replace them.

Adaptive round OpenAI executions Claude executions Total
1 12 11 23
2 95 75 170
3 34 27 61
4 10 14 24
5 8 10 18
6 1 2 3
7 1 0 1
8 1 0 1
Total 162 139 301

These are scenario executions, including repeated unresolved cases, not 301 unique scenarios. The latency measurements separately record first visible text including profile extraction and generation-only first text. Different case sets prevent round/model medians from being read as a speed ranking.

Three additional focused Claude executions are kept outside those counts and the 558-case benchmark:

The random C49 simulator had taken a passing sponsorship branch; that did not prove the earlier Miami ticket branch was fixed. Replaying the actual failing intent and turn order was necessary. The report distinguishes these focused executions from the benchmark result.

FAQ freshness and withdrawal tests

The original live Pinecone suite passed seven mechanical scenarios: add, edit, existing-chat refresh, long-to-short stale-tail removal, FAQ withdrawal, last-FAQ deletion and file deletion. Initial answer-update samples were roughly 9–11 seconds, with one existing-chat sample at 2.87 seconds. These measurements start from the test upload/update path, not an audited Sanity publish-to-answer SLA.

Mechanical deletion alone was insufficient: the assistant could invent a removed room detail or repeat a removed document's welcome phrase. The first prompt-only response fix failed and remains archived. The narrow current-source check then addressed the observed behavior.

Final application-response probes are 4/4 verified: two per model, covering removed FAQ and removed file facts. Both models returned inability to confirm plus Support, rather than stale/fabricated details. Probe durations were 899–1,393 ms. These narrow deterministic fallbacks do not need a model token stream.

Only unique initially empty test namespaces were changed. Removal was verified before querying the app. Final cleanup confirmed empty namespaces twice; an intermittent Claude namespace listing took 17 seconds to settle. No test enquiry was captured, and temporary environment copies were removed.

Mechanical suite · OpenAI responses · Claude responses · Earlier unsuccessful attempt.

Remaining content and operational checks

10 confirmed content-gap occurrences remain in the retained benchmark:

Missing approved source Occurrences Required content work
Volunteer roles, eligibility and application 6 Publish the current guidance/teams and refresh the existing event knowledge.
Complete sponsor names 2 Add approved company names or explicit alt labels to image-only logos, then refresh.
Miami speaker destination 1 Publish the intended speaker page if applicable. Current speakers/agenda-schedule pages are unpublished; do not invent a public link.
Decision-maker percentage 1 Publish an approved statistic. Attendance and Fortune 500 leadership figures do not establish it.

Content shares are 5/190 (2.6%) of original raw UAT, 5/171 (2.9%) of adjusted original UAT, 10/558 (1.8%) overall raw, 10/491 (2.0%) overall adjusted, and 10/185 (5.4%) of the reviewed remaining cohort. These percentages describe test occurrences, not missing pages or knowledge-base volume. Honest uncertainty can pass while the underlying requested fact remains unavailable.

The three retained operational occurrences are C47 on both models (independent browser/session identity isolation not exercised by a single-chat simulator) and C50 OpenAI (actual SMTP bounce handling untested; corrected-address summary and accepted 200 were observed).

Before release, separately verify real Zapier → Salesforce/Postmark delivery, bounce handling, browser/flyout lifecycle, live cookie/production/tool controls, configured Sanity publish-webhook filters, PDF extraction and sustained publish-to-answer timing. These are untested operations, not confirmed application defects. The local receiver and isolated source probes cannot certify them.

Verification, evidence and reproduction

Final local verification: 264 web library tests + 39 shared AI tests = 303 pass, 0 fail. Touched-source Biome checks are clean. The web typecheck still exits with existing errors elsewhere; the final scan found zero touched-file errors, not a clean whole-repository typecheck. GROQ ingestion scope was verified on real data in both published and raw perspectives.

Useful evidence:

Existing checks can be reproduced without sending external enquiries:

# Run each command from its owning package.
cd apps/web
bun test src/lib
cd packages/ai
bun test

Re-rendering the documentation/report reads retained artifacts and incurs no model calls. Instructions are in the report maintenance note. An intentional paid replay should preserve the manifest, same approved case contract and loopback-only submission receiver; do not restart whole batches merely to rebuild documentation.

Separate event-route repair

The earlier October 5 website route issue is separate from this local AI fix cycle. A missing/unpublished Miami agenda route could fall back to another event's page with the same short slug, exposing the Paris agenda under a Miami URL. Exact full-path matching now prevents that cross-event fallback.

Actual extract from eventRoute.ts:

export const EVENT_ROUTE_MATCH_GROQ =
    'defined(fullSlug) && fullSlug == $fullSlug'

Unlike the AI changes above, the saved deployment record shows this two-file repair deployed to main at 6ba30a04. Its isolated regressions recorded 11 pass, 0 fail, separately from the 303 AI/library checks.

The production smoke recorded six expected statuses: both unpublished Miami agenda paths return 404, the Paris agenda paths return 200, and both event roots return 200. The saved record also reports browser verification and retains a Miami 404 screenshot. This is earlier deployment evidence, not a new live check or deployment during this documentation task.

Technical architecture · Simple call guide · Tool guide · Facts and workflow policies · Editor operations · Launch controls · Privacy, cookies and submissions.