- Austin's Newsletter
- Posts
- Why AI Tools Lie to You
Why AI Tools Lie to You
AI thinks it's right, but you just steered it wrong

Your AI has been lying to you.
You ask it to do something and it says "Done" with complete confidence. But when you look closer, you notice it made something up, skipped something you told it to do.
It said the task was finished, but in reality…it wasn’t.
This is just a symptom of a bad system. A system that starts at the top with the AI CEO (you).
Last week we built the system that runs without you. This week is the harder part: How do we keep AI honest and make sure it’s doing what we want it to do?
It all comes down to one thing: properly monitoring your systems. And that requires getting real metrics, just like any CEO does.
So what does that look like in practice?
What Gets Measured Gets Managed
“What gets measured gets managed.” Corporate speak? Yes. But is it facts? Yes.
You would never let a new hire work for a month with zero check-ins and just assume it's going fine. You establish standard, measurable guidelines to ensure they’re on the right track.
There are three things that every good metric will have:
A number (countable, not just vibes)
A target (what "good" is, defined before the run)
A cadence (when you look at it)
For a new sales hire, that would be something like the number of accounts closed compared to the goal of 5, over the course of a month.
The moment you stop measuring, the output starts drifting. The new sales hire closes 4 accounts, then 3, and then your whole team misses the target for the quarter.
It works the same way with AI.
The automation you made keeps running, and the status says “green” or “working”. So you think it’s working, but the output slowly stops matching the reason you built it. You don't notice until it’s done major damage. I call this Green Light Drift.

This is already happening with some of the biggest companies in the world. Ford had to rehire 350 engineers after leaning too hard on AI for quality control. IBM tripled its entry-level hiring once it hit the limits of automation. Commonwealth Bank reversed the job cuts it made for an AI voice bot that underdelivered. Every one of them trusted AI more than they had checked it, and Green Light Drift took over.

The truth is - what gets measured gets managed, so you need to have metrics for your AI’s performance.
And a good metric needs to be simple and concise. If you’ve got something like an AI email sorter automation, here’s what it would look like:

So look at how you use AI, and any automations you’ve setup and answer: What is a quantifiable metric that this can communicate (i.e. Number of emails)
What is a target outcome for this (i.e. zero emails left unsorted since last run)
And what is the cadence (i.e. run every Friday at 8AM)

If you can't fill in all three, you won’t know if your automation is effective. You’re vibe building, which…is will cost you time and money…
The Failure Hides Until It's Expensive
I built an automated CRM to audit my inbound coaching emails and route them. It kept missing threads. I assumed it was an AI prompt problem, so I kept refining the prompt.
The prompt was fine. The MCP (the plug that connected Claude to Gmail) was wired to my Ops Manager Sage’s email, [email protected], instead of MY EMAIL. It was reading the wrong inbox and missing every thread that only landed in mine, which wasn’t obvious because Sage is CC’d on a lot of my emails.
And because I had no log of which account the tool actually hit, the failure was invisible. If I had one line to check which inbox it read, I would have caught it in a day instead of weeks.
I blamed AI. The real problem was the CEO. ME.
Anthropic's engineering team named this exact failure mode. In their guide on evaluating AI agents, they describe what happens when you wait for problems to surface on their own:
"Problems reach users before you know about them."
Whether you’re building for users or just for yourself, you have to catch this drift before it compounds. And that means logging the right things.
Make AI Leave a Note
After your automation runs, have it send you a specific line: what it did and one number that proves it actually did that. That's a Driver Log, and it will be driven by the Number, Target and Cadence we established earlier.
Not "ran successfully." That just means it turned on. You want "sorted 41 of 41 new emails" A number you can check in one glance.
It's the same standard you would hold a person to. I tell my team we need to communicate delivery timelines with an exact time; it’s not End of Day, it’s 5PM EST. One is explicit and verifiable. The other is vague.
If 5PM comes and I don’t see any output, I know instantly that something is wrong. Hold your AI agents to the same bar.

To help you set this up, open the automation you picked earlier and paste this in.

That's the fix I was missing in my CRM disaster. Now every automation I run has a log just like this. When something breaks, I don't spend weeks guessing at the prompt. I open one file and see exactly which run went wrong, and why.
But here’s the critical part. Not everything should be watched the same way.
New Hire, You Watch. Proven, You Delegate.
Every automation writes a Driver Log. The only thing that changes is who reads it.
A new automation should be treated like a new hire.
You watch it closely and make sure it’s doing what’s expected.
For your AI Automations, the best way to do this is run it locally, read the overview logs yourself and closely monitor. (My favorite way to do this is use Claude’s Desktop App and use their Routines feature).
Then once you’ve run it over 3 or 4 times successfully, you can move it to the cloud.
To do that - within Routines, you can select the cloud option, which will run it remotely.
But no matter what, you still have to keep a close eye on it…

You Are the Auditor
At the end of the day, you have to be auditing and driving the system. The monitoring strategies I mentioned here are only effective if you’re ACTUALLY reviewing them.
AI can EXECUTE, but it’s the JUDGMENT that steers. If you outsource both, the system stops belonging to you and will look like it’s working, but it will slowly drift and waste you money.
So this week, do two things:
Add one Driver Log. Pick one automation. Give it one success criterion and one line that proves it hit.
Block 15/20 minutes at the end of your week to review these logs. You can thank me later 🙂
Now if you got this far, you’ll like these 3 updates from my life
1. I finished the Ironman 2 days after my 30th birthday, 14 hours of swimming, biking and running 🤦
2. I am rolling out BuildPartner 2.0, you can learn about it here. It’s a custom Claude Plugin that will walk you through, step by step, the exact builds that I cover in my YT videos/Newsletters. This makes your life like 100x easier. It’s free and has had 3,563 signups to date.
So check that out, and I’ll SEE Y’ALL IN THE NEXT ONE.

Whenever you're ready, here’s how I can help:
For Business Owners & Executives:
1. Apply for my Executive AI Coaching Program: Linked here
2. Want to build a SaaS product without hiring a CTO? Linked here
For Everyone:
3. Use BuildPartner.ai to build faster with Claude Code (try free): Linked here
4. Forwarded this email? Sign up here