The Automation Log
How to evaluate AI claims from vendors (and founders)
Learn how to cut through AI vendor demos and founder pitches with the right questions — failure rates, human-intervention counts, and the edge cases that reveal the
Evaluating an AI vendor comes down to three moves: get off the demo, get to the failure data, and stress-test the edges. Vendors who build real systems welcome that process. Vendors who built a compelling slide deck will stall, deflect, or reschedule.
What does a real AI system look like versus a polished demo?
A real system runs on your inputs, not theirs. The fastest way to find out which one you’re looking at is to hand the vendor a messy, real-world scenario mid-meeting — something from your actual operations — and watch what happens. A rehearsed demo is always on rails. A live system handles the unexpected.
Ask to see the system in a state it wasn’t prepared for. Ask what the last three support tickets were about. Ask to watch a failure get handled in real time. If the answer is “we’d need to set that up” or “let me get back to you on that,” you have your answer.
This isn’t adversarial. It’s the same discipline I apply to treating systems like employees — you don’t evaluate a new hire on their best day. You evaluate them on a hard Tuesday.
How do I read failure rates and human-intervention counts?
Failure rate and human-intervention count are the two numbers that matter most, and most vendors don’t volunteer either. Failure rate tells you how often the system produces a wrong or incomplete output. Human-intervention count tells you how often a person has to step in to fix or approve something.
A system that claims to be fully autonomous but has a high human-intervention count isn’t autonomous — it’s a workflow with expensive human QA baked in and not disclosed on the pricing page. Ask for both numbers. If the vendor says they don’t track them, that’s a red flag. If they track them but won’t share them, that’s a bigger one.
For context on where humans genuinely belong in an automated workflow, see Human in the loop: where people belong in an automated business. Human checkpoints aren’t a flaw — undisclosed human checkpoints are.
| Question to ask | What a good answer looks like | Red flag |
|---|---|---|
| What is your failure rate? | Specific number with a definition of “failure” | “We don’t really track that” |
| How often does a human intervene? | Tracked metric, broken out by task type | “Rarely” with no data |
| What happens at the edge? | Live demo or documented failure modes | “Great question, let me follow up” |
| Who is liable when it’s wrong? | Clear contract language | Deflection to terms of service |
The edge cases that end the meeting
Edge cases are where real systems separate from demo systems. Pick the three scenarios from your operations that are messiest — ambiguous inputs, missing data, emotionally charged interactions, off-script requests. Feed them to the system. Watch what happens.
For voice AI specifically, I look at what happens when a caller goes off script, gives conflicting information, or asks something the agent wasn’t trained for. What AI voice agents can and cannot do (told straight) covers the honest boundaries. Any vendor who claims their voice agent handles everything without exception is either lying or hasn’t run it in production long enough to find the exceptions yet.
Edge case handling also reveals architecture. A well-built system fails gracefully — it routes to a human, flags for review, or says it doesn’t know. A poorly built system confidently produces wrong output. Confident wrongness is worse than an honest handoff.
Red flags that end the meeting
Some signals are disqualifying. I stop the conversation when I see:
- Benchmarks without methodology. “Our AI is 94% accurate” means nothing without a definition of accurate, a test set description, and a date.
- Refusal to show errors. Every system has errors. A vendor who says otherwise hasn’t been in production.
- Vague autonomy claims. “Fully automated” with a human review team in the background is a pricing model, not a product claim.
- Demo-only environments. If I can’t touch the live system, I’m not buying the live system.
- Founder who can’t explain what breaks it. The founder of a real AI product knows exactly where it fails. That knowledge is a sign of operational maturity, not weakness.
In the businesses I run — including an AI receptionist platform and a real-estate brand operated end-to-end on automation — the vendors and systems I trust most are the ones who lead with their failure modes, not their highlight reels. As of September 2026, the single most reliable signal of a production-grade AI system is a founder or vendor who can describe, without hesitation, the last three things that broke, what caused each failure, how the system handled it, and what was changed afterward. That level of operational honesty is rare. When I find it, I move fast. When I don’t, I walk.
How founders should think about this from the other side
If you’re a founder reading this before a meeting with an investor or buyer: the right move is to lead with your failure data. Walk in with a clear-eyed breakdown of where your system struggles, what your human-intervention rate is, and how you’re reducing it. That’s not weakness — it’s the signal that you’re actually running the thing in production.
I work with companies on this as a Fractional Chief Automation Officer — helping them build systems that hold up under scrutiny, not just under ideal conditions. The goal is a stack that an informed buyer can stress-test and still want to own.
For a broader framework on evaluating automation investments before capital moves, see Measuring automation ROI honestly (including the failures).
The platform I built for AI reception, Business Runner, went through exactly this kind of scrutiny during development. The hard questions made it better.
Want to put these questions to a live AI system right now? Try the voice agent on this site and see how it handles the edges.
Questions people ask
What questions should I ask an AI vendor before buying?
Ask to see the system live on an edge case, not a rehearsed demo. Request data on failure rates and how often a human has to intervene. Ask what happens when the AI is wrong and who is liable for the output.
What are red flags when evaluating AI vendor claims?
Red flags include benchmark numbers with no methodology, demos that run only on pre-loaded data, founders who deflect questions about errors, and contracts that bury unlimited human-review clauses. Any claim that the system is fully autonomous deserves deep skepticism.
How do I tell if an AI product is real or just a wrapper?
Ask the vendor to describe the model underneath, the fine-tuning or prompt layer, and the fallback when the model fails. A real product has a clear answer. A wrapper with no moat usually pivots to UI and integrations as the differentiator.