FAQ and uploaded-document freshness tests
Index
Cases
This operational block extends O03 (FAQ freshness) and O12 (document ingestion). It is separate from conversational UAT and uses distinctive synthetic facts rather than an LLM judge.
| ID | Change | Required evidence |
|---|---|---|
| FAQ-F01 | Add a FAQ | Its vector is readable, semantic search retrieves it, and a fresh visitor receives the new fact. |
| FAQ-F02 | Edit its answer under the same FAQ ID | New text replaces the previous answer in fetch/search and a new visitor response; an unrelated FAQ stays unchanged. |
| FAQ-F03 | Revisit it in the original conversation | The original chat ID and old answer remain in history; asking for the latest FAQ must return the updated fact without repeating the stale answer. |
| FAQ-F04 | Replace an uploaded FAQ file with a shorter version | Baseline creates multiple chunks and yields the original answer. Replacement yields the new answer and leaves exactly the new chunk; old tail chunks disappear. |
| FAQ-F05 | Remove one FAQ from the synced set | Full FAQ sync removes its vector, semantic retrieval no longer returns its answer, and a fresh chat does not repeat it. The unrelated FAQ remains. |
| FAQ-F06 | Remove the last FAQ | Empty FAQ sync removes the remaining FAQ, while the uploaded document remains intact. |
The tests use the application's FAQ formatting, FAQ sync, file upload, retrieval and chat paths. FAQ fixtures enter at the sync boundary; they do not edit the shared Sanity dataset. These are not ordinary run.ts chat cases: they require controlled source changes and stage-level evidence.
Timing
The runner records:
- Sync/upload duration, including extraction and embeddings.
- Time until the new vector text is observable.
- Time until semantic retrieval returns the new fact and excludes the old one.
- Time until the agent answers with that fact.
- Removal and stale-tail cleanup time.
All update measurements start immediately before the sync/upload call. Each vector/search stage requires two consecutive matching observations at two-second intervals. Chat probes run at fifteen-second intervals. Observed timings are upper bounds at this polling resolution, and later stages include the time spent checking earlier ones. The existing-conversation case separately measures question-to-answer after the update.
Default timeout is 180 seconds per stage, adjustable through FAQ_FRESHNESS_TIMEOUT_MS (maximum ten minutes). This is a diagnostic threshold, not an agreed production freshness SLA. A timeout is a failed test with elapsedMs: null, not a successful propagation measurement. Run multiple samples before presenting a typical or p90 result.
Run safely
Billing clearance is still pending for the paid retest. This runner has not been executed against live services yet. It makes embedding and conversation calls and must wait for that clearance too.
Use a dedicated local app with a new empty namespace, never the existing dev/staging/prod FAQ store. Both app and runner must use the same values. Keep the API secret in the existing private environment file.
From the repository root, start the test app (example namespace; choose a new unique suffix every run):
PINECONE_CONSOLIDATED_INDEX=unleash-knowledge \
PINECONE_INDEX_GENERAL=uat-faq-unique-run-id \
AI_MODEL_CONCIERGE=openai/gpt-5.4-mini \
AI_MODEL_AGENDA=openai/gpt-5.4-mini \
AI_RATE_LIMIT=10000 \
ZAPIER_LEAD_WEBHOOK_URL=http://127.0.0.1:8843/zapier \
pnpm --filter web dev --host 127.0.0.1 --port 3063
Run the test from apps/web, after billing clearance:
PINECONE_CONSOLIDATED_INDEX=unleash-knowledge \
PINECONE_INDEX_GENERAL=uat-faq-unique-run-id \
FAQ_FRESHNESS_BASE_URL=http://127.0.0.1:3063 \
bun --env-file=.env scripts/faq-freshness.ts --live
The runner verifies /api/pinecone/status resolves the actual app's general store to that exact namespace and that it is empty before writing. It rejects remote app URLs, shared FAQ namespace names and mismatched app/runner configuration. Cleanup removes only the new namespace's test records and records success/failure explicitly. Never use the example namespace if it already contains data.
Repeat with anthropic/claude-sonnet-4.6, a new namespace, and another isolated app port. Record application source hashes and model settings with the wider defect rerun. No Sanity settings need changing; no lead or email flows are requested. The loopback webhook prevents accidental downstream delivery if a workflow is misclassified.
Sanity publish timing
The automated block does not measure editor save-to-publish or publish webhook delivery. To cover the entire CMS path, run a separate O03 integration check in a disposable Sanity dataset with its own webhook and isolated Pinecone namespace:
- Record the FAQ publication acknowledgement timestamp and document revision.
- Record receipt/start/finish of the signed
/api/pinecone/syncwebhook. - Record the same vector/search/answer stages above.
- Change the answer and repeat, then unpublish/remove it and confirm withdrawal.
- Verify a saved draft alone never changes published knowledge.
- Keep a webhook-disabled/failed-delivery control separate; measure recovery through manual refresh or the scheduled refresh rather than counting it as immediate publish propagation.
Do not use the shared production dataset for synthetic facts. This CMS publication check remains untested until the disposable dataset and webhook are configured. The current application FAQ publish handler performs a full FAQ resync; nightly refresh is a fallback. Uploaded files use the upload route rather than the FAQ-set publication webhook.
Reading the results
The runner writes results.json, models.json, cleanup.json and a short report.md under docs/reports/faq-freshness/<run>/. Every probe preserves elapsed time and source/response evidence. Publication-to-answer latency must not be inferred from sync-to-answer measurements.
A failed answer after successful retrieval points toward tool selection, stale conversation context or response generation. A stale vector points toward ingestion/upsert; fresh fetch but stale search points toward index visibility/search. Missing old tail removal points toward replacement cleanup. A webhook that never arrives is a separate CMS/infrastructure issue. Keep these categories separate in the compiled report.