Model-independent checks

Index

Index

Package tests

Ran once against the same frozen application source used by both model variants. Selecting a response model cannot affect these pure-code tests.

Package Passed Failed Evidence
ai 34 0 log
config 1 0 log
form-tests 89 5 log
forms 56 0 log
forms-migration 23 0 log
sanity 6 0 log
seo 10 0 log
ui 20 0 log
web 212 0 log
webgl 7 0 log
Total 458 5

The five failures are in formBuilderConditions.test.ts: the fixture calls new HTMLInputElement() with happy-dom preloaded, which raises TypeError: Illegal constructor. An isolated rerun reproduces the same failures. These are test-harness failures; the full ordinary suite is not green. Isolated evidence.

Build and typecheck

Biome: pass across all 25 included changed source files. Log.

Studio production build: pass. An initial missing nested media-plugin dependency was restored from the existing workspace installation; no application source was changed. Build log.

Web project typecheck: fail (exit 2). None of the reported errors matched the 25 included working-tree source changes. This is not a clean repository typecheck. Log, changed-file scan.

Operational checks

Case Status Evidence / remaining scope
O01 Response latency Measured locally Per-turn profile extraction plus first-token latency in paired report. Local concurrent load, not standard-broadband production certification.
O02 Inbox delivery Not exercised Local receiver acceptance only; no real email sent.
O03 FAQ freshness Not exercised Would require a shared CMS content mutation.
O04 Dynamic content freshness Not exercised Reads live sources, but no timed publish/update experiment.
O05 Tool/source configuration Partial audit Paris maps event-world/event-world-agenda plus Paris-filtered sitemap; Miami maps event-america/event-america-agenda plus Miami-filtered sitemap. Consolidated index uses legacy store name as namespace. All namespace contents not exhaustively audited.
O06 Log PII redaction Fail in sampled tool trace Synthetic name/email stored unredacted in submitLead arguments; evidence.
O07 Salesforce correctness Not exercised downstream Capture payloads can be inspected, but no real CRM records were created.
O08 Session identity Partial Distinct generated test IDs; browser conversation survived reload/navigation. Parallel real-visitor isolation not independently proven.
O09 Rate limiting Unit coverage only No sustained abusive external traffic generated.
O10 Widget rendering Sampled desktop/mobile Cookie gating, hotel link and agenda flyout checked; tight mobile padding noted. Evidence.
O11 Kill switch Unit coverage only Global/shared Sanity settings not changed.
O12 Ingest integrity Not exercised No upload/clear/replacement against shared stores.

Disabled-workflow, cancellation, price precision, profile-memory and consent cases are included in web/AI unit tests. Legacy inactive Jev classification experiments are excluded from this active-system model comparison; current selector and workflow acceptance blocks are included.

October 5 targeted correction and retest

The original results above remain the baseline. After fixing the native-input fixture, the affected form package passed 94/94 tests. Changed web/AI packages passed 222/222 and 35/35, including new regression checks. The complete ordinary suite was not rerun.

The targeted defect retest completed the 37 selected prior failures and 14 further runs of still-failing cases after additional fixes. Passing cases and obsolete expectations are excluded. Six isolated FAQ freshness tests are prepared but have not run; no update propagation time has been established.