
What Happened When Our 37-Agent AI Team Learned to Watch Itself
In one working day, we rebuilt how our 37-agent AI team operates. Not the producing side. The writing, research, and reporting were already humming. What we added was the layer almost nobody builds: agents that watch the other agents, a written compass they all aim at, and a rhythm that turns every repeated mistake into a permanent fix. This is the case study of that day, with real numbers, and the three moves a boutique business owner can copy without touching anything technical.
If "AI team" is a new phrase for you, start with our earlier post on what an agentic team is and how to build one. This post picks up where that one ends. You have the team. Now who manages it?
The discovery that forced the upgrade
Two findings made the old way impossible to ignore.
First, the same bookkeeping defect hit our team roster four separate times. Every fix was honest, and every fix was local, so nobody ever saw the repeat. Four smart corrections, zero learning.
Second, over a 30-day stretch, six pieces of work quietly went stale inside the system, and every one was caught by the human founder. The owner had become the quality system of last resort, which is exactly backwards. If you are personally the one noticing that your AI keeps making the same mistake, you are doing a job an agent should hold.
The theme underneath both: most AI teams are built to produce. Almost nobody builds the layer that watches, aims, and improves the producers. That layer is where compounding lives.
Upgrade one: hire a watchdog before another producer
The first new hire was not another writer or researcher. It was a watchdog: one agent whose only job is reading the system's own records weekly, then reporting what keeps recurring, where two agents duplicate work, and which agent keeps failing for the same reason. Every finding arrives with a proposed fix, written as a concrete edit ready for approval.
Its first memo proposed ten fixes. Tracey went through them by number the same night and approved most on the spot.
The design choice that makes this safe is diagnose-only. The watchdog reads everything and changes nothing without a yes. That bookkeeping defect from earlier? A watchdog reading the logs would have caught it on occurrence two, not four.
Upgrade two: a goals file before wider autonomy
Autonomy without a compass produces busy agents and random output. So before widening anyone's leash, we wrote one file naming the north star, the active bets in priority order, and the constraints. The rule the file installs: any agent proposing new work ties the proposal to a named bet or says plainly that it fits none.
The constraints half turned out to be as load-bearing as the goals half. Our file states that money always waits for a human yes, and that certain pricing never gets quoted externally. That means ambitious agents stay ambitious inside the fence.
Upgrade three: shorter instructions won, three to nothing
Here is the counterintuitive one. We tested whether our agents' long, procedure-heavy instruction files were helping or hurting. We took three writer skills and produced each brief twice, by separate agents that did not know they were in a test, once with the full instructions and once with a slim version cut 53 to 60 percent. A blind judge picked the slim output three times out of three.
The pattern in the verdicts was telling. Losers lost on one small bent rule each, and the rule-benders were always the procedure-heavy versions. The guardrails carried the discipline. The scaffolding diluted it.
One rule kept the experiment safe: never slim the quality gates. Identity, brand facts, and hard limits stay. Micromanagement goes.
Upgrade four: events, not just schedules
A weekly review means a failed payment can sit unnoticed for six days. So we designed the pipeline: business events, a failed payment, a new lead, a cancellation, post to one channel the minute they happen. Twice a day, a sweep turns each alert into a drafted response, timed just before the windows when approvals happen anyway. The event feed is a small one-time build inside the CRM, and ours is going in now. Real-time eyes, same-day hands.
On top of that sits one Monday briefing. Every report, every project's open loops, and every waiting decision folded into a single post under 600 words, decisions first, oldest first, each showing its age in days. That last detail matters most. A decision that has quietly waited three weeks stops being quiet.
The housekeeping that paid for the whole day
Three small rules came out of the cleanup, and one of them likely saved a client relationship.
Delete the twin. Whenever anything gets replaced, an invoice, a draft, a page, the same session deletes the old copy or writes down who deletes it and when. One sweep found five live orphans, including a duplicate recurring invoice that could have double-billed a client. Old versions left alive can act.
Keep the failures. Rejected drafts and flopped posts stay on file as first-class records, because the failures are where the watchdog finds the patterns.
Every number ships with its context. No naked metrics in any report. Each number carries a versus-prior-period line, and outcomes lead volume. The question a report answers is never "what is the number." It is "is the number good."
What a boutique business owner should copy first
You do not need 37 agents. You need the watching-and-aiming layer, and it starts with three moves.
- Write the goals file. One page: your north star, your bets in order, your hard constraints. If you cannot name a revenue target yet, write the file anyway and mark that line as the open question. Agents aiming at named bets beat agents guessing.
- Schedule a weekly self-review. Have your AI read its own recent work, your notes, and your corrections, then report what happened twice or more. Repeats are system problems wearing an instance costume.
- Consolidate to one briefing. Decisions on top, oldest first, with age in days. One reader beats fewer reports.
And underneath all of it, keep the constitution that makes the autonomy safe: the AI drafts, the human sends. Money waits for a yes. Nobody grades their own work.
This is the system we run our own agency on, invoice sweep receipts and all, and the same one we build for boutique businesses.
FAQ
Do I need a big AI team for any of this to matter?
No. The producing layer scales with your business, but the watching layer works at any size, even one owner and one AI assistant. A goals file, a weekly pattern review, and a single decisions-first briefing require no code and no headcount.
What keeps an AI team this independent from doing something wrong?
Three rules that never move. The AI drafts and the human sends, so nothing reaches a client without an owner's eyes. Money always waits for a human yes. And no agent ever checks its own work. Autonomy went up at our agency precisely because those fences stayed put.
How does a self-improving AI team connect to being found on Google and in AI search?
Volume without discipline drifts, and drifting content stops earning citations. A watchdog keeps quality consistent as output grows, which is what search engines and AI answer engines reward. The structural side of that discipline, how content gets shaped so machines can quote it, is covered in our 5 Gears of Visibility framework.
What if I can't define my business goals clearly enough for a goals file?
Write it anyway and mark the fuzzy lines as open questions. The file's job is direction, not perfection. Even a draft compass stops your AI from proposing work that serves no bet at all, and your weekly review will sharpen it month by month.
Ready to see what this looks like in your business?
We build AI teams like this for boutique and luxury businesses across Marin and the Bay Area, watchdog, goals file, and all. If you want a digital presence that improves itself while you run your company, book a call with Lens on Luxury and we will map your first three moves together.
