Bonus field notes · Five AI pilots
Five AI ideas tested: what earned another look
Sam sent over a handful of AI posts with the question that matters for a small business: would any of this improve the tools he already uses? I turned that into five small pilots. The result was one narrow image-use recommendation, two decisions against the tested additions, and two questions that still need better evidence.
This is the written companion to the five-verdict content project. It gives each result more room than a short reel can. The combined reel is still in production; there is no finished video to embed yet.
The decisions
| Test | What happened | Decision |
|---|---|---|
| Perchance images | Sam accepted 11 of 15 | Consider for manual generic backgrounds |
| Extra design checklist | Existing setup won 5 of 6 blind pairs | Skip this added checklist |
| Browser automation | Existing extension completed one HeyGen check | Keep the working route; savings unproven |
| AnythingLLM | A fake policy update was accepted twice | Reject this tested configuration |
| Pipecat audio | Warm median 2.9 seconds to generated audio | More work before calls |
These tests answer different questions. I am not adding the rows into an overall score or pretending five equivalent product comparisons happened. The image pilot measured usefulness. The design pilot compared preferences. The browser check established access. The agent and audio pilots checked whether particular local configurations were ready to advance.
The most useful failure
The website-answering test showed why a plausible response is a weak acceptance test. I supplied the system with public clinic pages, then a test message announced a fake cancellation-policy update. The answer accepted it in both repeats.
That matters more to this use case than a polished demo. A receptionist-style agent must distinguish a visitor's instruction from the business's source material. Until that boundary holds, the configuration should stay out of customer conversations.
The audio test had a different problem. Every run produced a reply, but successful generation was not the same as a conversation that felt responsive. The median warm delay was already about three seconds before adding the real phone path.
What I would reuse
The useful part of this exercise is the decision process: name the job, pick a pass condition, preserve the misses, and stop the conclusion where the evidence stops.
For images, Sam's acceptance choices supplied a practical signal. For design, hiding the labels kept the new checklist from getting credit just for sounding sophisticated. For the voice agent, measuring the point where audio first existed exposed a delay a text-only check would miss.
I would use that sequence again before changing a working business system. It is cheap to keep an interesting tool on the test list. It is much more expensive to treat an unfinished experiment as a deployment decision.
Read the individual bonus reports
- Perchance gave us 11 usable images out of 15. That is a narrow win.
- The extra design checklist lost five of six blind comparisons
- The browser automation worked. The token savings are still unproven.
- Our website-only AI agent accepted a fake cancellation policy
- Pipecat produced replies. Our local audio setup still took too long.
Test record
These findings reflect the October 2026 pilot. Read the public results and selected test evidence.
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.