How Do You Know If Your Agency's AI Is Actually Working?

Share
How Do You Know If Your Agency's AI Is Actually Working?
Most agencies can't answer this question. They subscribed to the tools, ran the pilot, maybe even built a few automations — but when someone asks "what's different?", the room goes quiet. The test is simple: can you name one decision your AI made better this week, with a number attached? If you can't, it's not working. It's just running.

Everyone uses AI now. Almost nobody can prove it does anything. That gap has nothing to do with the tools. It's about whether anyone's measuring what they actually do.

I run an agency. We run on an AI system we built — the same one we install for clients. I don't say that to flex. I say it because I've been on both sides of this gap, and the difference between "using AI" and "AI is actually working" is smaller than you'd think. It comes down to one habit.

Can You Name a Decision AI Made Better This Week?

This is the only question that matters, and most agency owners can't answer it.

Not "did you use AI this week" — you probably did. You prompted something, ran a tool, maybe generated some copy. That's activity. The question is whether AI changed a decision you would have made differently without it.

Here's what that looks like from our ad account: we launched a creative, and within 40 impressions — not 40 clicks, 40 impressions — the system flagged it for a $3.02 kill. Same-day decision. Zero clicks, zero wasted spend beyond the test. A human media buyer would have let it run for days "gathering data."

That's a decision AI made better. I can point to the number. Can you?

What Happens When Nobody's Measuring?

The scary version of "AI isn't working" isn't when it fails loudly. It's when everything looks fine and it's quietly broken underneath.

We had a nine-day stretch where our email delivery system had a single missing character in one line of code. Dashboards were green. Reports looked normal. But buyer emails weren't arriving. Nine days of revenue leaking through a crack that every metric said didn't exist.

We caught it because a buyer messaged asking where his product was. After that, we built a layer that checks whether the thing the system sent actually arrived — not whether the platform said it sent it. That's the difference between using a tool and owning a machine: the machine checks its own work. But only after you've felt the pain of not having it.

Most agencies running AI don't have that layer. They see the dashboard, it says "sent," and they move on. The gap between "sent" and "arrived" is where the money disappears, and nobody notices because nobody's measuring at the right depth.

What Does AI "Working" Look Like When It's Actually Failing?

Here's one that burns agencies who think they're using AI well.

We tested six "AI-improved" variations of an ad creative against the original. Better copy. Sharper hooks. Tighter CTAs. The AI did exactly what we asked. All six died — $41.64 burned across 373 impressions with zero landing page views. Meanwhile, the original — the ugly, unpolished one — closed a sale the same day.

If we weren't tracking at the individual creative level, we would have seen "ad set performing" and never known that every AI-generated variation was dragging the original down. That's AI "working" by every surface metric and failing by the only one that matters: did someone buy?

AI wrote better copy than I did. All six versions died. The lesson: faster production without measurement just means you produce more things you can't evaluate.

How Do You Actually Tell If It's Working?

Forget the dashboards for a second. Here's the diagnostic I use:

1. Can you point to a kill? Not a win — a kill. Something AI told you to stop, that you stopped, and you can see the money you saved. We identified that a specific placement was hemorrhaging budget across every campaign, killed it permanently, and the savings compounded every day after. That's a kill with a number — about $300 before we caught it.

2. Can you show compounding? Is this week's output smarter than last week's because the system learned something? Our retargeting audience was 20 people. Twenty. It generated an $8.88 cost per acquisition — on a product where we could afford to spend over sixty dollars. That wasn't luck. It was a system that had been reading every buyer interaction for months, building an asset no competitor could copy. The data compounds. The AI without the data is just software.

3. Can you point to a catch? Something the system flagged that a human would have missed. The nine-day email gap. The placement bleed. The creative variation trap. If your AI has never caught something you didn't see, it's not monitoring — it's just producing.

If you can't answer yes to at least one of these, your AI is running but it isn't working. You've adopted it. You haven't implemented it.

What's the Real Divide in Agencies Right Now?

It's not between agencies that use AI and agencies that don't — almost everyone uses it now. The divide is between agencies that have built the feedback loop and agencies that are just running tools with no way to know if they're producing anything.

The agencies building a machine underneath their business? One that reads its own outputs, catches its own failures, gets smarter every week on the client's data? You can't undercut those guys. Not because they use AI, but because they own something that compounds. Cancel the subscription and the data, the patterns, the intelligence — they stay.

The ones renting tools and checking dashboards? Their clients are already doing the math on whether the $5K retainer is worth more than a subscription that does the same surface-level work.

You don't prove AI is working by showing someone a tool. You prove it by pointing to decisions, kills, catches, and compounding — with numbers attached.

If you want to see what that system actually looks like — the one that compounds on your clients' data and makes you impossible to replace — we broke down the entire $25/day approach here.

Frequently Asked Questions

How long does it take to know if AI is working in your agency?

You should be able to point to a measurable decision within the first 30 days. If after a month you can't name a single kill, catch, or compounding pattern with a number attached, the implementation is superficial. The tool might be running, but the feedback loop isn't built.

Is it enough to just track revenue to know AI is working?

Revenue is a lagging indicator — by the time it shows up (or doesn't), weeks of spend are already gone. The leading indicators are kills (what did you stop?), catches (what did the system flag?), and decision speed (how fast did you act on data?). Revenue follows those.

What's the most common sign that AI isn't actually working?

The clearest signal is when nobody can point to a decision that changed because of AI. The team uses the tools. The dashboards look fine. But nothing about the actual workflow or decision-making is different from before the tools existed. Activity without measurement is the most expensive kind of AI adoption.

Can you use AI effectively without building your own system?

You can use AI tools effectively for individual tasks — generating copy, summarizing data, drafting emails. But the compounding effect that makes AI a business advantage requires a system that connects those tasks, tracks their outcomes, and feeds results back into the next cycle. That's the difference between renting AI and owning it. Here's what owning it actually looks like.

What's the first thing to measure when implementing AI?

Start with time-to-kill: how fast can your system identify something that isn't working and stop spending on it? If you can get that under 48 hours for ad creative and under 24 hours for delivery failures, you're ahead of most agencies. Our fastest was $3.02 and 40 impressions — same-day.


I document how a real agency actually runs on an AI system — real campaigns, real spend, real numbers, updated as it happens.

Get it by email →