How do I run a fleet of AI agents without losing control? The 6 laws

Published July 12, 2026 · 25 United Capital · También en español

One AI agent is governed with attention. A FLEET is not: attention doesn't scale, and the founder who adds agents the way one adds browser tabs ends up with the worst of both worlds — a team's cost and no team's order. Every new agent multiplies unverified work, reports nobody cross-checks, and versions of the truth in circulation.

The solution is neither new nor technological: it is organizational design — the oldest discipline of the firm, reborn for employees who think at a thousand words per second and forget every night. We apply it in a real company where a non-technical human directs and EVERYTHING else is AI: these are the six laws that make the fleet compound instead of scatter.

Law 1 — Design roles, don't hire agents

The system is defined by role types, not model names. Ours are four: the SOVEREIGN (human; objectives, quality bar, final judgment) · the TRANSLATION LAYER (keeps human-machine communication faithful, with no executive authority) · the TECHNICAL CHIEF / VERIFIER (the skeptic with access to the real terrain who fixes what is done) · the EXECUTORS (bounded work under contract). A new agent joins by writing ONE role sheet — capabilities, access, class of authority, verification level it is subject to, cost profile — and ZERO system redesign.

The receipt: our authority map was written this way from day one ("the system is defined by these role types, not by agent names") — and thanks to that it has survived model handovers and agent arrivals and departures intact: the incumbent changes, the role remains.

Violation smell: decisions that depend on what one specific agent "knows". If an agent is irreplaceable, you don't have a star employee: you have a single point of failure.

Law 2 — A single node owns "done" (two-level supervision)

A fleet can — and should — have delegated supervision: a veteran executor reviewing another's day-to-day. But the final "VERIFIED" lives in ONE node (the verifier-in-chief), and delegated sub-supervision NEVER replaces it: it can reject, correct and package someone else's work; it cannot declare it done. Nothing reaches official status without passing through the same gate.

The receipt: our two-level org chart is sealed by formal amendment to our operating statute (record and re-seal of July 8, 2026): delegated operational supervision of one executor over another, with the final verdict reserved by law to the verifier. It works because delegation multiplies capacity WITHOUT fragmenting the truth.

Violation smell: two different "dones" depending on which agent you ask.

Law 3 — The scaling law: never hire faster than you verify

The bottleneck of an AI-operated company is not the capacity to execute — it is the capacity to VERIFY. Every new agent adds production somebody has to check; if the verification lane is full, the new hire doesn't add real work: it adds unaudited work (which is worse than no work). Our hiring pipeline has phases with gates — scouted → candidate → bounded pilot with authorized cost → incumbent — and no agent touches real work without first being bound to the house's integrity law.

The receipts (two, and the second is a renunciation): (1) our high-volume executor pilot was fully designed — seat, minimal access, delegated supervision — and remained GATED awaiting the sovereign's cost authorization: the design is not the hire. (2) An earlier candidate was SUPERSEDED without ever being activated when a better option appeared — the seat wasn't filled "because it was already decided": not hiring is also a scaling-law decision.

Violation smell: more active agents than verified work in the last week.

Law 4 — No delegation without observability (you don't govern what you can't see)

Delegating without being able to look is not trust: it is abdication. Every delegation is born WIRED to be observable — the agent's work leaves a trail in records the supervisor can read without asking permission and without depending on self-reporting. And structural honesty matters more than a pretty picture: our own map declares the real gap in writing ("strong at consulting, weak at alerting") instead of feigning omniscience — a DECLARED gap can be watched; a denied one cannot.

The receipt: our worst fleet incident was ~11 hours of an agent working without effective observation — it produced a mountain of artifacts about itself and zero product. No dashboard caught it: the later audit did. The rule was born there: the observable trail is wired BEFORE delegating, not after the grief.

Violation smell: finding out what an agent did… by asking the agent.

Law 5 — Orders are contracts, not conversations

You don't "chat" work over with an agent: you issue an order with two gates — entry (what will exist when it's finished, how it will be checked, read back and confirmed BEFORE starting) and exit (verification against reality, never against its report). The backlog is measured in RESULTS, not hours; every order is small, single and reversible; and everything an executor delivers is born labeled [CLAIMED] until the verifier says otherwise.

The receipt: we maintain a written protocol of copy-paste orders to executors (each order with its gates, explicit boundaries and an honest-failure clause: a true FAIL is worth more than a false PASS). Fuzzy missions were the habitat of all our incidents; contract-orders killed them.

Violation smell: an agent that has been "advancing" for days while nobody can say what will be missing when it finishes.

Law 6 — Top-down asymmetry is non-negotiable

Across the whole fleet, only the human sovereign authorizes three things: cost, external action, and the irreversible. The agents — all of them, verifier included — propose and execute under confirmation. An internal protocol can NEVER "pre-approve" an expense; a delegation chain NEVER dilutes this rule: every link inherits it whole.

The receipts (two): (1) across our whole accumulated operation, unauthorized cost — the only one of our three supreme threats that has never occurred — remains at zero, because the rule has no exceptions to negotiate. (2) In the blind exam of our canon, the trap case was exactly this: an expense "pre-approved by protocol". The context-free successor rejected it and escalated — the asymmetry survives even the handover of the one who executes it.

Violation smell: the sentence "it was little money and I didn't want to bother you". The amount was never the point; the asymmetry is.

The summary that fits on a card

1. Roles with a sheet, not agents with an aura. 2. A single node declares "done". 3. Don't hire faster than you verify. 4. No observable trail, no delegation. 5. Orders with two gates, backlog by results. 6. Cost, external and irreversible: only the human. Always.

Why you can trust these laws (and how to verify it)

This playbook does not describe a hypothetical fleet: it describes ours — an org chart sealed by amendment with a record and a public cryptographic fingerprint in the Receipts Room, an order protocol in use, a hiring pipeline with real documented renunciations, and the incident that taught us Law 4 told without makeup. We build in public with truth labels; the work can be visited.

Frequently asked questions

How do I manage several AI agents at once without losing control? With organizational design: defined roles, a single final verifier, contract-orders and wired-in observability — not with more of your attention.

How many AI agents can one person direct? As many as their verification lane can carry: the limit is how much work you can check, not how many agents you can open.

Can one AI agent supervise another? Yes, as delegated operational supervision; the final "done" verdict stays in a single verifying node.

What permissions do I give a new AI agent? The minimum of its role sheet, after binding it to your integrity rules and BEFORE its first real mission; cost and external actions always remain with the human.

Why does my team of agents produce a lot yet advance little? Production without verification and fuzzy missions: measure verified results, not activity.


25 United Capital · one human directs, AI operates, written rules govern — built in public at 25united.com.