Bonus field notes · Five AI pilots
Our website-only AI agent accepted a fake cancellation policy
A test message claimed the clinic's cancellation policy had changed to none. The website-answering agent accepted it, even though the retrieved material contained the real policy. The same result appeared in both repeats.
That is enough for me to reject this configuration for customer-facing use. A visitor should not be able to rewrite a business policy by declaring an update inside the conversation.
The setup
The pilot used AnythingLLM 1.17.0 with a local Qwen model identified as qwen3-8b-ctx12k, a frozen set of 35 public website pages, query mode, six retrieved results and a similarity threshold of 0.25. Reasoning output was disabled through the native stream route. Each question began in a fresh thread.
Thirty questions ran twice, producing sixty answers. All sixty completed, and each question's second answer matched its first. That gives repeated observations in this setup, not sixty independent cases or a universal AnythingLLM accuracy score.
No private client records or live booking actions were part of this evaluation.
The failures worth acting on
The policy-injection case was the clearest boundary failure. The input presented an instruction as a supposed system update, and the answer treated it as authoritative despite the retrieved policy.
Another answer retrieved conflicting cancellation percentages but chose one without acknowledging the disagreement. A staff question missed the relevant material and then overstated the result as an absence of names on the website. A separate answer emitted a citation on example.com that was not a valid source URL.
I also had to correct the evaluation key. The price of $110 did appear in the frozen welcome blog. It would have been wrong to call that number fabricated. The citation was the error in that case. Checking the evaluator against the source is part of the test, too.
What this says about website-only agents
Loading website pages into a retrieval system does not, by itself, prove that answers will stay within those pages. This pilot exposed several separate requirements: retrieve the right passage, handle conflicting passages, preserve the distinction between user instructions and source authority, and attach a real citation.
I would test those behaviors separately before moving a configuration into a voice agent. An answer that sounds natural can still fail every useful business rule around it.
The decision and next test
The verdict is no to this configuration. It is not a claim that every AnythingLLM deployment fails, and I did not run a comparison showing the current clinic agent outperformed it.
The next attempt should first resolve conflicting source content, then test instruction handling and citation validation using the same frozen questions. Only after the text behavior passes would I advance to source-refresh tests, multi-turn conversations and voice. This pilot did not advance to those stages.
Test record
These findings reflect the October 2026 pilot. Read the public results and selected test evidence. Return to all five bonus reports.
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.