0

The operating system behind my 31 AI agents

Technical cut of Managing 31 AI Employees (2026-04-05).

Since Lunar New Year I have been running a digital team on OpenClaw, with Cursor Ultra, Claude Code Max, and several coding plans. The team now has 31 agents—"four departments and one office": Discovery, Product, Growth, a Strategy think tank, and a CEO-office secretary. Forty-four scheduled tasks fire every day: morning brief, evening report, content review, dashboards, health checks.

On March 8 I handed daily operation of Global Tech Events to the agents. A month later it was still collecting, filtering, organizing, and publishing with no intervention from me. I also built an Agent Town pixel prototype for multi-agent collaboration, and in a Demand Discovery experiment the agents produced a usable version in a little over a day—most of my time went into infrastructure, not the build.

Harness is necessary; organization is the missing layer

Silicon Valley's answer for reliable agents is Harness Engineering. OpenAI's Codex experiment is the clearest public case: three engineers, roughly one million lines of code, five months, about 3.5 PRs per engineer per day, every line written by agents. They made design decisions readable in Markdown, enforced a layered dependency direction (Types → Config → Repo → Service → Runtime → UI), used ARCHITECTURE.md and linters as feedforward and feedback, and ran garbage-collection agents against drift.

That solves one agent plus one task chain. Past thirty agents, I hit organization-level problems a single harness does not cover: who gets which slice of context, how agents hand work to each other without me, and how results survive past one conversation.

On March 31, 2026, a Claude Code source-map leak (59.8MB, about 510k lines of TypeScript) showed production answers I adopted: do not store what can be re-derived; stale memory is worse than none; a background process (autoDream) and, in my system, a knowledge agent review the base; subagents work in isolated worktrees; the lead agent must still understand intent before delegating.

The hourglass

A human org is a pyramid. Mine is an hourglass. Above me: unlimited tasks. Below me: nearly unlimited execution. The bottleneck is my attention. The first topology was a star—every agent talked to me. Past ~30 agents that collapsed. I became the hourglass neck, not the tip of a pyramid. The fix was not "manage harder." It was protocols.

Seven-step startup read

Every agent boots with a mandatory sequence:

  1. Read COMPANY-STATE.md—a 50-line live board (stage, today's priorities, pending signals, blockers), updated nightly by the secretary.
  2. Read its role definition.
  3. Read the knowledge-base protocol (where to find and write, not the whole corpus).
  4. Read my identity profile and annual plan.
  5. Scan its signal inbox.
  6. Drop expired signals.
  7. Begin work.

Named roles in practice: CMO D'Addario, CTO Sweeping Monk, Demand Discovery Sherlock Holmes, secretary Liu Yifei. The same fact lands differently in each role's slice. Loading everything into every agent produced long, bland output. The seven-step read is the hallway conversation, made explicit.

Signal files

Cross-department coordination is not a meeting. It is a signal file: type, priority, expiry, recipient, and an SLA. Example: a high-priority demand must be consumed by the CPO within 24 hours; a metric anomaly must escalate to me within 4 hours. Past SLA, the secretary marks it yellow in the 8am brief; at 2× SLA it turns red.

When the development pipeline finishes, a pipeline-done signal goes to the CMO and growth so go-to-market can start on the next boot. Agents stay async. The shared knowledge base keeps them consistent.

Four-layer memory

I keep the knowledge base in four layers:

  • Context — who I am, values, annual plan, current focus. Stable; loaded at every start.
  • Inbox — unprocessed flows: seeds, notes, signals. Short-term memory.
  • Kernel — axioms and mental models. Agents cannot change these without my authorization.
  • Library — domain knowledge, industry notes, contacts, tool instructions. Long-term memory.

Every write needs YAML metadata (date, author, type, tags). No metadata means the document does not exist for search. Before writing, a three-step check: Can this be derived from existing materials? If yes, do not store it. Does it have a shelf life? Market data defaults to 90 days and needs valid_until. Is it a fact or a judgment? Facts rot; the why of a decision lasts.

Claude Code's rule—"stale memory is more dangerous than no memory"—replaced my earlier habit of saving everything.

Three operating rules

Standardization. The CMO runs a self-review checklist: concrete opening (person or scene), no forbidden words, open ending instead of a neat summary, sounds like a person when read aloud. If length, vocabulary, platform format, and style score all pass, I do not review it. I only see failures.

Two-way doors (Bezos). Type 1 decisions are irreversible—new platform positioning, major brand-voice shifts, language that could create controversy. I decide. Type 2 decisions are reversible—a weak Xiaohongshu post, a suboptimal publish time. Agents decide. Most decisions are Type 2.

Async structured deliverables. Output is a self-contained document: what was done, why, evidence, uncertainties that need me. I batch-review on my schedule. Chat-bound output breaks async.

Possession (and who manages me)

After an important deliverable, every critical agent must: give a judgment, not only data; offer 2–4 next-step options; then execute the handoff (write to the knowledge base, notify downstream, create the signal). Do not leave the ball on the floor for me to pick up.

Thursday evening is Memory Garden: Buffett, D'Addario, and Li Ziqi review my week from business, content, and practice angles. The challenge directive requires them to flag when my actions conflict with written values, when obligation is driving a decision instead of curiosity, when a plan violates "simple over complex," or when frameworks are covering avoidance. Format: quote the values line, then name the deviation.

What failed

In early April I disabled two scheduled tasks after repeated failures. The star topology past ~30 agents failed for a clearer reason: I was the bottleneck. The stack I use now is:

Digital Organization Design → multi-agent coordination, information flow, knowledge, decisions
Harness Engineering          → one agent's tools, constraints, feedback, acceptance
Context Engineering          → context for one invocation
Prompt Engineering           → instruction for one invocation

The useful work is not cheering for more agents. It is writing the protocols that keep them from consuming the only scarce resource in the system—my judgment.


Source: darren-su.com


All rights reserved

Viblo
Hãy đăng ký một tài khoản Viblo để nhận được nhiều bài viết thú vị hơn.
Đăng kí