Report maintenance and evidence

Index

Scope

This is the completed October 5 failure-only review. It is not a new complete execution of all 279 cases per model. See the report, rendered report and implementation fix record.

Application changes remain local and unpushed. Test receivers were loopback-only; no real enquiries were sent. The test apps are stopped and FAQ test namespaces were cleaned. Do not restart model batches or alter deployed configuration to rebuild documentation.

Rebuild documentation only

Run each command from the repository root. They read retained test artifacts and write local report/derived JSON files; they make no Gateway or other network calls.

python3 docs/reports/2026-10-05-ai-defect-retest/remaining-fix-cycle/compile_remaining.py
python3 docs/reports/2026-10-05-ai-defect-retest/remaining-fix-cycle/sync_report.py
node docs/reports/2026-10-05-ai-defect-retest/remaining-fix-cycle/render_reports.cjs

Edit report prose in compile_remaining.py, then rebuild. sync_report.py adjusts relative links for the parent report and preserves its #results anchor. HTML is rendered from the respective Markdown source with the installed Marked dependency. Do not modify a generated raw verdict to make the percentages better.

Artifact contracts

Artifact Role / restriction
Original full-model-comparison results Baseline: 279 case occurrences per model; supplied UAT is C01–C95 only.
Earlier parent/round2–4 results Earlier failure-only replacements, retained separately from this cycle.
remaining-register.json Defines the 185-occurrence reviewed cohort.
round1–round8 transcripts/results 301 scenario executions including repeated unresolved cases. Round 1 includes combined-results-only evidence for some cases; loading individual result files alone loses real results.
review-ledger.json / .csv Transcript/source review, without replacing raw verdicts.
retained-results.json Latest result per model/block/case, 279 per model, with evidence location.
retained-scores.json Current raw/adjusted before/after scores and all 67 evidence-only exclusions. Content and operational checks remain scored.
Parent adjusted-scores.json Historical 33-exclusion input. Preserve it; do not replace it with the current 67-exclusion output.
quote-recovery, quote-recovery-2, quote-information Three separate focused executions: first failed quote, corrected passing quote, passing information turn. Outside the 301 execution and 558 retained-result counts.
faq-response-retest Four final verified withdrawal responses and cleanup. Separate from the seven original mechanical freshness scenarios.
faq-response-retest-attempt1 Retained unsuccessful prompt-only attempt. Never silently replace it with passing evidence.
quote-information/manifest.json Final tested source: base commit plus SHA256-listed files. Documentation edits do not change that application snapshot.
verification.json, cleanup.json Final local checks and test resource cleanup.
Parent report-before-remaining-cycle.* Historical compiled snapshot; keep it unchanged.

Only explicit captured 2xx responseStatus values count as accepted handoffs. A 503 or an older capture without status is not independently verified acceptance. No local capture establishes downstream email or CRM delivery. An adaptive simulator can choose a different branch; a passing random branch does not prove an earlier failure was reproduced.

Application verification

The retained final evidence records 264 web-library tests and 39 shared AI tests passing, clean touched-source lint, and no touched-file TypeScript diagnostics. The whole web typecheck has existing failures elsewhere. Documentation rebuilds do not rerun these tests.

Any intentional future paid replay needs an explicit case selection, frozen app/configuration, the current approved behavioral contract and loopback-only handoffs. Keep earlier evidence and source refreshes identifiable. A fresh full-suite result must be labeled as a new run, not merged into these historical counts without an explanation.