UNLEASH AI assistant: technical architecture
Contents
- 1. System overview
- 2. One message, end to end
- 3. Event context and session memory
- 4. Jev: narrow workflow selection
- 5. Workflow behavior
- 6. Tool inventory
- 7. Data authority and freshness
- 8. Prompts and models
- 9. Consent, side effects, and trust boundaries
- 10. Streaming and observability
- 11. Configuration and operational controls
- 12. Verification evidence and remaining gaps
- 13. Where to change what
- 14. Response policies
- 15. October fix record and reports
Documentation updated: 5 October 2026. The original September deployment snapshot was 91cec88a on pre-prod. This guide now incorporates the locally tested October changes on base b2534f60 plus the final source manifest. The October fixes are not pushed or deployed. Published settings were not changed during this fix cycle; the requested refresh of five published Paris pages into Pinecone was applied. This is an implementation guide, not a live configuration audit or a guarantee that every answer is correct.
The assistant combines a streaming conversational model, a small workflow selector, deterministic application logic, structured data lookups, semantic retrieval, and contact-request tools. Jev selects the appropriate existing workflow. It does not generate visitor-facing answers, validate those answers, decide consent, or send enquiries.
1. System overview
flowchart TD
UI[SolidJS chat widget] --> PROFILE[POST /api/ai/profile]
PROFILE --> UI
UI --> API[POST /api/ai]
API --> GUARDS[Validate input, resolve event, apply guards]
GUARDS --> STATE[Read active workflow from process memory]
STATE --> SELECT[Continue directly or ask Jev for destination]
SELECT --> FLOW[Existing workflow rules and prompt assembly]
FLOW --> DIRECT[Deterministic reply when applicable]
FLOW --> MODEL[Streaming concierge through AI Gateway]
MODEL <--> TOOLS[Scoped tools]
TOOLS <--> SOURCES[Sanity, Pinecone, Inwink]
TOOLS --> HANDOFF[Validated Zapier handoff]
DIRECT --> DELIVERY[Response delivery]
MODEL --> DELIVERY
DELIVERY --> UI
API --> TRACE[Langfuse telemetry]
DELIVERY --> REVIEW[Asynchronous conversation review]
REVIEW --> TRACE
The two main packages are apps/web (SolidStart API routes and application services) and packages/ai (the SolidJS widget, shared conversation types, event/consent helpers, and retrieval ingestion utilities). Configuration defaults live in packages/config; editorial configuration and content live in Sanity.
External responsibilities:
| Service | Responsibility |
|---|---|
| Vercel / SolidStart | Host the web application and server routes; run scheduled refreshes and background completion work. |
| Vercel AI Gateway | Access the concierge, profile extraction, Jev, and review models. |
| Sanity | Published pages, FAQs, media, prompt/model configuration, badge content, and ticket-pricing mappings. |
| Pinecone | Semantic knowledge retrieval and stored agenda records. |
| OpenAI embeddings API | Embed source chunks and search queries for Pinecone, separately from conversational Gateway calls. |
| Inwink | Event/agenda source integration and current public ticket-product pricing. |
| Zapier | Accept lead and agenda-email requests for downstream processing. |
| Langfuse / OpenTelemetry | Trace model/tool activity and attach conversation-review scores. |
2. One message, end to end
- Collect browser context. The widget keeps a chat ID, history, extracted visitor profile, event context, and conversational consent. It supplies the current page path plus browser language/platform/timezone hints. These hints are not verified company, role, location, or eligibility facts.
- Extract explicit profile updates. The widget awaits
POST /api/ai/profilebefore requesting the answer. A typed model extraction uses the visitor's message, the previous assistant question, existing profile, and page context. Failure retains the existing profile and transcript. - Submit the chat request.
POST /api/aireceivesmessage,type,location,chat_id,history,userProfile,selectedEvent,consent, and browser context. Supported surfaces includechat,global,agenda,sales-agent, andsales-agent-spex. - Validate and constrain input. The server applies input limits, rate/duplicate-message checks, and injection-pattern checks. It resolves the event and applies relevant cancellation and deterministic workflow rules.
- Choose the destination. With the new selector enabled, the server reads active workflow state. Routine active-flow replies skip Jev; other turns use an entry or continue/change decision. Uncertain/error results fall back to existing behavior.
- Run the existing workflow. Deterministic branches can answer directly. Otherwise the server assembles prompts, known-profile context, event scope, and available tools, then calls
streamTextthrough AI Gateway. - Execute tools. The conversational model chooses among allowed tools, except where existing orchestration forces an appropriate handoff step. Tool implementations enforce their own input, event, consent, and delivery checks.
- Deliver the response. Ordinary text streams incrementally. Authoritative agenda and handoff results can replace generated wording. A known send request is held until the tool outcome is available, so that path can report success or failure accurately.
- Finish background work. Stream completion/abort finalization flushes telemetry. A separate response monitor schedules conversation review after delivery. Some early deterministic responses follow a shorter path.
Current limits include a 4,000-character incoming message, at most 20 history turns / 80,000 history characters, 3,500 output tokens, and eight model/tool steps. The widget has a 60-second inactivity watchdog and abort/request-sequence handling to prevent stale responses from overwriting a newer interaction.
Sources: chat route, widget, profile route, delivery.
Supporting HTTP endpoints
POST /api/ai: streamed chat response.POST /api/ai/profile: typed profile extraction and merge.GET /api/ai/models: effective concierge/agenda/profile model IDs and configuration sources, used by Studio.GET /api/ai/agenda?eventType=…: structured source agenda and Markdown export; this is not the personalizedbuildAgendatool./api/pinecone/*: upload, catalogue refresh, and administrative retrieval operations; mutation handlers require internal authorization.GET /api/cron/pinecone-refresh: authorized scheduled full refresh.
3. Event context and session memory
Event resolution
The implemented precedence is:
- An unambiguous event named in the latest message.
- Previously selected conversational event context.
- Event inferred from the current page path.
- The most recent unambiguous event mention in user history.
Paris / UNLEASH WORLD map to paris; Miami / UNLEASH AMERICA map to miami. Legacy path aliases are supported. If the event remains unknown, event-dependent tools ask for clarification instead of silently selecting a show. Browser timezone is not an event selector.
Conversation and workflow memory
| Memory | Storage and purpose | Important limit |
|---|---|---|
| Browser conversation | sessionStorage: chat ID, transcript, visitor profile, chat-scoped consent. |
Browser-session state, not authenticated account identity or cross-device memory. |
| Active workflow | Bounded in-process map keyed by chat_id, containing workflow, event, and update time. |
It is not durable across serverless instances or restarts. |
Active workflows are tickets, spex, or agenda. Information is a turn destination, not a fourth persistent commercial workflow. Event mismatch or expiry means there is no applicable active workflow. Successful handoffs and cancellation clear it. A rejected handoff ends qualification and receives an honest Support fallback; the visitor must explicitly request a new attempt. Routine acknowledgements cannot silently restart collection or submission.
The in-process map is capped at 5,000 entries and expires entries after one hour. It contains no contact details and never represents consent or send authorization.
The profile is rendered as a known/missing values table in the model prompt. Explicit name, work email, company, role, interests, and ticket details should be reused rather than repeatedly requested. Extraction does not infer consent. Existing facts are merged; new booking/event intent can reset ticket-specific details. This improves context availability but does not guarantee the conversational model will never re-ask a known field.
Sources: event and consent helpers, process-local workflow state, profile parsing and formatting.
4. Jev: narrow workflow selection
The deployed selector is controlled by AI_JEV_FLOW_SELECTOR=active. It uses typesafe-ai/jev through AI Gateway and experimental_evaluate to choose a typed destination.
| Situation | Behavior |
|---|---|
| No active workflow | Classify into tickets, SPEX, agenda, information, or uncertain. |
| Active workflow + routine short reply | Continue without calling Jev. Examples include recognized affirmations, help, an email, or a quantity. |
| Active workflow + another message | Ask the narrower continue/change question. Explicit new tasks can switch flows. |
| Information side question | Answer the question while preserving the active workflow for later continuation. |
| Low confidence, invalid output, timeout, or failure | Fall back to existing workflow behavior rather than invent a decision. |
Jev receives the current message, page path, active workflow, and up to six prior turns (roughly three exchanges). Oversized historical turns are shortened. The selected probability must meet a threshold: 0.65 for information and 0.75 for other destinations. Evaluation has a 1,500 ms timeout and no retries. These thresholds are application policy, not proof that probabilities are calibrated.
The result guides the existing route and narrows tools: non-agenda destinations lose agenda construction/email tools; agenda and information destinations lose submitLead. Information turns bypass ticket qualification and forced commercial handoff. Fallback behavior is deliberately less restrictive.
Boundary: Jev chooses where the request belongs. Existing code and the conversational model still handle qualification, field correction, pass recommendation, consent, confirmation, and tool execution. Jev does not review or rewrite streamed answers.
The older AI_JEV_ROUTING_MODE and AI_TICKET_JEV_MODE paths remain separate switches and are off in the staging configuration described here. Earlier broader experimental action-routing results should not be attributed to this narrower implementation.
Sources: decision policy, Jev call, integration.
5. Workflow behavior
Tickets
Explicit booking/help/contact intent enters the existing qualification flow; asking for prices or pass options alone remains informational. The assistant reuses known event, quantity, company/role and contact details. A requested group quote can proceed with no pass selected; quantity ranges remain estimates until explicitly corrected. Pass facts and current numeric prices come from badgesLookup, with deterministic ticket-answer helpers for supported questions. Repeated choice-help advances to a recommendation. Missing or conflicting facts stay unavailable rather than becoming invented prices.
The final action is a delegate enquiry, not payment, ticket issuance, or a confirmed booking. A successful submitLead means the webhook accepted the request. Registration links can also be supplied from maintained source content.
SPEX: sponsorship and exhibiting
The assistant uses commercial knowledge and event pages to explain opportunities, gather the visitor's requirements, and prepare an enquiry. Saved personal details can be reused, but the intended behavior still requires a current summary and affirmative send confirmation. A previously saved consent decision is not a universal instruction to send future requests.
Cancellation/withdrawal is recognized before routing and stops the pending contact flow. The cancellation history boundary prevents an older affirmative reply from being treated as fresh authorization. Cancellation cannot retract a request already accepted downstream.
Agenda
agendaLookup handles complete-list factual questions. buildAgenda creates personalized schedules from structured sessions instead of asking the model to improvise a timetable.
The planner excludes incomplete or restricted sessions, scores topic matches, avoids overlapping sessions, allows a ten-minute transfer between different rooms, and selects at most six sessions per day. Matching is lexical; it is not a learned preference optimizer. Alternatives are returned separately.
The server owns the plan's formatted reply. Session URLs use the event's agenda-schedule route with agendaEvent and agendaSession, supporting the site's session flyout behavior. Emailing a plan validates selected URLs against the event's available sessions and formats the payload server-side.
Selected-session recovery still uses recognizable formatted agenda evidence in the transcript, followed by server validation; it is not an authenticated durable plan store. The October cycle fixed and retested requested-day/source-year handling, profile reuse, corrected recipients and agenda-email follow-ups. Browsing selections are validated into a plan before email handoff. See sections 12 and 15 for final evidence and operational limits.
Information and content
The assistant answers event, hotel, venue, article, webinar, and general questions using appropriate sources. Latest articles and future webinars use structured date-aware queries. Hotel recommendations should expose the maintained external booking link, not merely a partner name or an internal travel-page link. pageDirectory provides labeled external links, and deterministic hotel handling covers supported requests.
Sources: workflow helpers, handoff helpers, ticket answers, agenda planner, agenda URLs.
6. Tool inventory
Tools are registered in the main chat route. Availability depends on page/event scope, surface type, consent, and selector destination; the model does not always receive all 18 registered tools.
| Tool | Purpose and source | Key inputs / constraints |
|---|---|---|
unleashAmerica |
Miami event knowledge, general knowledge, and relevant event pages. | Semantic query; event-scoped retrieval. |
unleashAmericaAgenda |
Miami agenda/FAQ semantic discovery. | query; use structured lookup for exhaustive questions. |
unleashWorld |
Paris event knowledge, general knowledge, and relevant event pages. | Semantic query; legacy World name. |
unleashWorldAgenda |
Paris agenda/FAQ semantic discovery. | query; not an exhaustive session list. |
generalKnowledge |
Brand/general FAQs plus current event material. | query. |
eventFacts |
Complete current-event/general FAQ entries and existing indexed event-page bodies for narrow policies. | Resolved event and supported topic; read-only; preserve conditions and published links. |
unleashSales |
Commercial knowledge store. | query. |
unleashSalesSpex |
Sponsorship/exhibitor knowledge store. | query. |
mediaCatalogue |
Semantic discovery of podcasts, reports, videos, and archive content. | query; not authoritative for newest articles or future webinars. |
siteMap |
Semantic page discovery with URLs and indexed event-page text. | query. |
agendaLookup |
Complete stored event agenda, exact sessions, stages, and days. | Optional event, stage, day, keynote, speaker, title, topic filters. |
badgesLookup |
Published badge pages and verified structured pricing. | Optional event; never treat prose/SEO snippets as numeric price authority. |
pageDirectory |
Published event-page directory with titles, URLs, headings, descriptions, and external booking links. | Optional event. |
recentArticles |
Direct Sanity query, newest publication first. | Topic; empty topic means latest overall. |
upcomingWebinars |
Direct Sanity query for confirmed future webinars. | No parameters. Empty means none currently listed. |
buildAgenda |
Deterministic personalized schedule over complete event sessions. | Topics, optional event/day; no invented sessions. |
submitLead |
Send a validated commercial/contact enquiry to Zapier. | Name, email, enquiry, allowed intent; optional company/role. Event and consent required. |
sendAgendaByEmail |
Submit an agenda-email request to Zapier. | Name, email, selected session URLs, optional event/opt-in. Consent and valid event sessions required. |
Semantic retrieval normally returns six results (agenda search uses fourteen), over-fetches candidates, limits repeated chunks from one page, filters known refusal-template content, and returns useful source fields such as text, title, URL, and question. Retrieval failures return an error/empty result, not a license to invent facts.
The new read-only eventFacts uses existing Pinecone namespaces rather than a different retrieval source. A narrow same-entity/same-field check also prevents unsupported answers for recognized named room/document facts after withdrawal. This is not a universal fact validator. Dedicated sales-agent and sales-agent-spex surfaces expose a narrower commercial tool set. Main event pages combine event-specific and shared tools. Handoff tools are consent-gated and retain checks inside their execution functions.
Sources: tool definitions and availability, retrieval filtering.
7. Data authority and freshness
| Question | Preferred authority | Freshness / limitation |
|---|---|---|
| Pass name, eligibility, inclusions | Published Sanity badge/attend content. | Depends on editorial accuracy and maintained mappings. |
| Numeric ticket price | Shared structured pricing resolver: Sanity mapping plus Inwink public products. | Configuration/product caches are 60 seconds; date bands are re-evaluated on read. Not instantaneous. |
| Explorer price | Application-only policy. | Say “Price on application”; do not quote a numeric amount from history, registration text, or SEO. |
| Session times, rooms, speakers | Complete agenda records stored in Pinecone. | Ten-minute warm-instance cache plus ingestion freshness. Complete relative to the loaded catalogue, not an independently live Inwink query each turn. |
| Latest articles | Direct date-ordered Sanity query. | Does not rely on similarity ranking for recency. |
| Upcoming webinars | Direct future-date Sanity query. | An empty result is different from archival webinar availability. |
| External hotel booking | Maintained event-page external links. | Missing/wrong source links still require content maintenance. |
| Broad FAQs / commercial knowledge | Pinecone stores populated from published content and uploads. | Dependent on refresh/upload success. |
Pricing reads published Sanity configuration without the CDN, maps passes to Inwink product categories, and handles date bands, tax metadata, and explicit overrides. Inwink requests have an eight-second timeout and bounded pagination. Conflicting/unavailable price data fails closed rather than using an unverified number. The curated live-pricing mapping currently covers Paris; equivalent Miami coverage must not be assumed.
Knowledge ingestion supports manual uploads and catalogue refreshes for FAQs, media, sitemap/event pages, and agenda. Publish revalidation provides incremental updates; a configured nightly cron at 03:00 UTC performs a full refresh. A configured schedule is not evidence that the latest run succeeded.
Source text is chunked and embedded using OpenAI text-embedding-3-small (1,536 dimensions) through the direct embeddings API with OPENAI_API_KEY. Records retain source metadata for retrieval and links. Query text is embedded for similarity search; structured agenda reads bypass embedding and fetch records directly. This is retrieval-augmented generation, not fine-tuning the conversational model on UNLEASH content.
PINECONE_CONSOLIDATED_INDEX switches logical stores into namespaces in one physical index; without it, the logical stores use separate indexes. Structured agenda lookup lists/fetches stored records and parses the full session set, avoiding top-K omissions.
Mutation endpoints require INTERNAL_API_SECRET as a bearer token; without a configured secret they are disabled. The scheduled refresh accepts CRON_SECRET or the internal secret.
Sources: badge loading, pricing server, pricing policy, ingestion and agenda loading, nightly refresh, publish revalidation.
8. Prompts and models
Sanity settings.ai supplies editorial prompts, additional instructions, purpose-specific ticket/SPEX/agenda consent notices, and task-specific model settings. The server caches this configuration for 60 seconds and combines it with application-owned workflow/handoff instructions and context.
| Task | Built-in default / fixed model | Notes |
|---|---|---|
| Concierge | openai/gpt-5.4-mini |
Generates streamed answers and tool calls. |
| Profile extraction | openai/gpt-5.6-luna |
Typed extraction of explicit visitor facts; 500 output-token limit. |
| Agenda model configuration | openai/gpt-5.4-mini |
Used for agenda workflow model turns; the buildAgenda planner itself is deterministic. |
| Workflow selector | typesafe-ai/jev |
Typed destination decision only. |
| Conversation review | openai/gpt-5.4 |
Separate post-response evaluator; override with AI_REVIEW_MODEL. |
For configurable concierge/profile/agenda tasks, precedence is task-specific environment override → Sanity model setting → built-in default. Concierge also supports legacy AI_GATEWAY_MODEL after the Sanity setting and before the default. These are source defaults, not a claim about the effective model of every deployed request; inspect resolved runtime configuration/traces for that.
Prompt rules instruct the assistant to reuse known facts, preserve event identity, use authoritative links/prices, ask for missing information, and summarize before sending. Those instructions are probabilistic behavior guidance. They are not equivalent to a deterministic validation check.
Sources: configuration assembly and resolution, default prompts/models.
9. Consent, side effects, and trust boundaries
Conversational contact consent is kept separately from extracted profile facts. A malformed email can be corrected without discarding an existing consent decision; withdrawal clears that decision. Email escape normalization handles text such as an escaped @ before validation.
Cookie/privacy preference dismissal is a separate UI concern. The privacy preferences overlay has a close action without accepting. Closing it does not grant conversational contact consent, and contact consent must not be conflated with tracking-cookie preferences.
Before dispatch, handoff code checks event, consent, required fields, email validity, and allowed intent. Server-known profile fields are merged into payloads. Agenda requests additionally validate session URLs and build their agenda text from server-owned records. Intent values distinguish delegate, agenda, brochure, startup, sponsorship, and registration-interest requests, with event suffixes such as UNW/UNA.
Both handoff tools POST JSON to ZAPIER_LEAD_WEBHOOK_URL with a 15-second timeout. A 2xx response means accepted by the webhook. Salesforce processing and agenda email delivery happen downstream; the application does not independently confirm CRM persistence or inbox delivery. HTTP 503 is a rejected attempt. Missing webhook configuration is a send failure, including on localhost.
There is no application-level guarantee of exactly-once downstream delivery. Any Zapier-side deduplication is a separate operational dependency. Saved profile/history/consent supplied by a browser are not an authenticated identity or a tamper-proof consent ledger. Current-request summaries and bounded natural confirmation have been fixed and retested, including a failed-then-passing exact Miami quote replay; these checks do not make the whole conversation an infallible transactional state machine. Lead payloads include opt_in: false; agenda opt-in is separate and defaults false. Agenda payloads include server-formatted Markdown and HTML.
Sources: payload and transport, handoff checks, consent recognition, privacy overlay.
10. Streaming and observability
Streaming
The server consumes AI SDK fullStream through a TransformStream and forwards text deltas as plain UTF-8 response bytes. The browser decodes and appends them to the assistant bubble. This is streamed text, not a client-facing structured tool-event protocol.
Authoritative agenda output and send results may be emitted as a complete server-owned message. For recognized handoff requests, generated prose is withheld until the outcome is known; if no send occurs, the response reports that instead. Consequently not every reply arrives token by token, even though ordinary model responses still stream.
The Markdown link transform normalizes quoted URLs and removes unsafe/empty destinations while holding only the current link, capped at 4 KiB. Model step boundaries preserve readable paragraphs. Jev is before the answer, so it can add time to first text. Profile extraction is also awaited before chat starts. The review evaluator observes delivery and runs afterward; it is not a synchronous answer-validation gate.
Traces and review
OpenTelemetry, the Langfuse span processor, and the AI SDK integration record model/tool work, including concierge, profile extraction, and Jev selection. Jev telemetry includes decision probabilities and usage. Finalization uses background work to flush traces.
When conversation review is enabled (default unless AI_CONVERSATION_REVIEW=off), a response monitor forwards chunks while capturing up to 24,000 characters. It records truncation/incomplete-stream evidence, then schedules a typed model review with a 30-second timeout and no retries.
Reviews classify topic and status, and can flag repeated questions, ignored withdrawal, wrong event, missing booking links, unsupported success, source contradictions, or unhelpful answers. Deterministic malformed-link checks are distinguished from potentially bad behavior identified by the model. Scores include conversation-review-status, conversation-topic, and conversation-confirmed-defect.
“No detected issue” is not proof of correctness. Incomplete evidence and evaluator unavailability are separate outcomes. The reviewer is a general model, not Jev, and never blocks the visitor's streamed answer.
Telemetry masking redacts many non-UNLEASH email addresses and phone patterns. It does not comprehensively anonymize names or all personal data. Browser session storage and webhook payloads contain contact details by design; retention/access controls need their own operational policy.
Sources: stream delivery, telemetry, response monitor, review evaluator, review schema.
11. Configuration and operational controls
| Setting | Function |
|---|---|
AI_GATEWAY_API_KEY / legacy AI_GATEWAY |
Gateway credentials. |
OPENAI_API_KEY |
Direct embeddings API credentials for retrieval/ingestion. |
AI_MODEL_CONCIERGE, AI_MODEL_PROFILE, AI_MODEL_AGENDA |
Task model overrides. |
AI_JEV_FLOW_SELECTOR |
active enables the new narrow selector. |
AI_JEV_ROUTING_MODE, AI_TICKET_JEV_MODE |
Independent older routing/interpretation experiments; off in this staging rollout. |
ZAPIER_LEAD_WEBHOOK_URL |
Contact/agenda request delivery endpoint. |
LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_BASE_URL |
Observability connection. |
AI_CONVERSATION_REVIEW, AI_REVIEW_MODEL |
Review enablement and evaluator selection. |
PINECONE_CONSOLIDATED_INDEX |
Optional shared physical index using logical namespaces. |
INTERNAL_API_SECRET, CRON_SECRET |
Administrative refresh authorization. |
AI_RATE_LIMIT, AI_RATE_WINDOW_MS |
Per-instance request throttling. |
Underlying Sanity, Pinecone, and Inwink client credentials are also required by their integrations. This table intentionally contains names, not secret values.
Pre-model protection defaults to 20 requests per 60 seconds per IP in an in-memory, bounded store. The third identical normalized message is blocked. Profile and chat requests use the guard, so a user-visible interaction can involve more than one guarded request. These controls are per instance, not a distributed abuse-prevention system, and can produce misleading duplicate warnings in shared-IP test runs.
The 18 September staging rollout enabled the narrow selector on Preview / pre-prod, keeping the older Jev switches off. To disable the selector, remove/change its active value and deploy that configuration; existing workflow behavior remains available. Production configuration is outside this rollout's scope.
Useful troubleshooting sequence:
- Correlate the visitor's
chat_id, event, transcript, and Langfuse trace. - Check active workflow, selector source/destination, and whether the turn bypassed classification.
- Inspect the exact tools exposed/called and their returned sources/errors.
- For stale facts, distinguish source content, ingestion/cache freshness, and model interpretation.
- For missing sends, inspect summary/confirmation recognition, consent, required fields, tool execution, and webhook status separately.
- For latency, separate profile extraction, selector time, source/tool time, first body chunk, total stream time, and browser rendering.
Source: guards.
12. Verification evidence and remaining gaps
The completed October failure-only review covers 185 previously non-passing model/block/case occurrences, with 177 actually retested across eight adaptive rounds (301 scenario executions, including repeats). Three focused quote/information executions and four FAQ-withdrawal response probes are separate. Latest cohort raw verdicts remain 138 pass, 27 partial, 18 fail, 2 blocked. Transcript/source review found 138 observed resolutions, 34 obsolete/evaluator-only judgments, 10 content gaps and 3 operational checks; no verified application defect remains open in that reviewed cohort.
The original supplied UAT remains a separate block: C01–C95, 95 cases per model. Combined raw retained pass rate is 48.4% before → 85.8% latest; adjusted using the same itemized exclusions is 53.8% → 95.3%. Across the entire 558-occurrence retained benchmark, raw pass rate is 57.2% → 85.7%, adjusted 65.0% → 97.4%. These are rolling results combining targeted reruns and untouched original results, not a fresh full-suite execution after the final change. Raw verdicts are never relabeled as passes.
Final local verification recorded 264 web library + 39 shared AI tests passing, clean touched-source Biome checks and no touched-file TypeScript errors. The whole web typecheck still has existing errors elsewhere. The separate live Pinecone suite passed seven mechanical update/delete scenarios; four final response probes verified honest uncertainty for removed FAQ/file facts on both models. Observed timings are samples, not a publish-to-answer SLA.
Remaining content is volunteer guidance, complete sponsor names, the unpublished Miami speaker destination and an approved decision-maker percentage: 10/558 (1.8%) retained occurrences. Browser/new-session identity isolation, SMTP bounces, real CRM/email delivery, live launch/cookie controls, configured Sanity publish hooks, PDF extraction and sustained source-update timing need operational verification. The local test receiver cannot certify them.
The earlier September selector comparison (18/23 raw passes enabled versus 13/23 disabled) and staging smoke remain historical evidence. They are not a comparison of the current October changes or proof of statistical improvement.
Evidence: current compiled report, complete fix record, review ledger, retained scores and exclusions, historical selector report, September deployment smoke.
13. Where to change what
| Change | Primary location |
|---|---|
| Widget behavior, transcript persistence, profile preflight | packages/ai/components/AiModule.tsx |
| Event resolution, consent/cancellation recognition | packages/ai/components/workflows.ts |
| Request orchestration, tool schemas and availability | apps/web/src/routes/api/ai/index.ts |
| Destination classification and confidence policy | apps/web/src/lib/flowSelector.ts, flowSelector.server.ts |
| Process-local workflow state | apps/web/src/lib/aiFlowSessions.ts |
| Prompt composition and model resolution | apps/web/src/lib/aiPrompts.ts, packages/config/ai-prompts.ts |
| Profile extraction and known-facts table | apps/web/src/routes/api/ai/profile.ts, apps/web/src/lib/userProfile.ts |
| Ticket answers, badge facts and pricing | apps/web/src/lib/ticketAnswer.server.ts, badges.ts, livePricing.ts, livePricing.server.ts |
| Agenda planning and selected-plan handoff | apps/web/src/lib/agendaPlan.ts, aiHandoff.ts |
| Webhook payloads and transport | apps/web/src/lib/leads.ts |
| Streaming and authoritative replies | apps/web/src/lib/chatDelivery.ts |
| Retrieval ingestion and catalogue refresh | packages/ai/components/api/pineconeUpload.ts, apps/web/src/routes/api/cron/pinecone-refresh.ts |
| Trace collection and post-response review | apps/web/src/lib/aiTelemetry.ts, aiResponseMonitor.server.ts, aiConversationReview.server.ts |
| End-to-end acceptance harness and fixtures | scripts/ai-acceptance/ |
Keep routing, workflow execution, data authority, and response review as separate responsibilities. A better destination decision cannot repair incorrect source content or a broken confirmation/send handler; each needs its own evidence and tests.
14. Response policies
Visitor-policy implementation is in aiResponsePolicy.ts. Explicit historical-edition requests receive a fixed redirect to current/future events; personal-information requests receive a privacy/data-processing referral to Support. Additional semantic phrasing is covered by application-owned instructions that override older CMS/source policy.
Standard contact addresses and the legacy name UNLEASH WORLD are normalized by a streaming text transform before response monitoring. Only a possible partial token is held between chunks; the full answer is not buffered. Direct contact questions receive one address: Customer Success for clear sponsorship/exhibiting, Support otherwise. These policies do not add a Jev answer-validation step or alter handoff consent requirements.
Policy test report records the local checks and their limitations.
15. October fix record and reports
The October fix record documents each observed defect, resulting behavior, code/test links, current source contracts and exact code extracts. It also separates original supplied UAT, rolling benchmark percentages, content gaps, focused failed/passing replays and operational checks. The current HTML report contains raw case verdicts and transcript evidence; the earlier report remains archived.