Field guide / sub-domain

AI Agents: Orchestration

39

sources in this field

Updated September 20, 2026

Current thesis

The shortest path to orientation.

This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.

Coordinator agents excel at batching work into 3–5 parallel streams with cross-model review (max 5 rounds) and adversarial verification. Production systems require versioned memory, concurrency control, five-stage loops with objective pass/fail verification. Task-level routing yields ~30% savings; heavy front-loaded planning prevents downstream chaos. Cold outbound GTM decomposes into four persistent agents (list builder, copywriter, campaign manager, infrastructure manager), each owning one artifact. Reliable multi-agent handoffs rest on three mechanisms: naming as join key, task-state-as-trigger (completion fires webhook waking next bot), 'blocking out loud' (agents post missing input and stop rather than guess). Claude Code as orchestrator with tools like ColdIQ MCP cut campaign builds from two days to one prompt and ~20 minutes. Ars Umbris functions as meta-harness, connecting multiple agent harnesses via MCP adapters to one durable workspace state. Parallelizing adversarial and browser-driven test runs across many instances makes exhaustive pre-release stress testing economically viable even for small teams, eliminating the tradeoff between thoroughness and resource constraints.

Evidence board

Claims worth carrying forward
01

A browser-based adversarial testing suite runs massively parallel adversarial checks against each new release, attempting to break it before ship, at a cost described as 'pennies.'

02

Parallelizing adversarial/browser-driven test runs across many instances makes exhaustive pre-release stress testing economically viable even for small teams.

03

Ars Umbris treats knowledge bases like codebases: typed files with explicit relationships and conventions an agent can parse, rather than freeform notes, so structure compounds across sessions instead of decaying.

04

A repo bundles six parts as local files: knowledge (notes/sources), types (definitions), skills (agent instructions), MCP tools (operations), projections (UI components), and agent instructions/profiles—all composable and git-trackable.

05

Repos declare dependencies via a manifest (repo.yaml with a deps list), similar to package imports in code, letting one repo (e.g. au-tree-research) supply vocabulary, skills, and MCP tools that other repos reuse without duplication.

06

Cross-repo references use explicit syntax: [[note::another-repo]] for wikilinks and source::au-base-types* for shared types, making dependencies explicit like imports—so a workspace of multiple repos reads as one typed graph.

07

Core thesis: knowledge work needs its own type engine (analogous to compilers/linters in software) because agents can silently accumulate broken links, missing fields, and inconsistent structures across sessions—reminding the model of conventions isn't sufficient at scale; independent checks are required.

08

Type definitions (YAML) specify required fields and valid reference targets (e.g., a claim type requiring a source::au-base-types field); when an agent writes a note missing that field, the engine flags a diagnostic the agent can read via tools and fix—though judging whether a source actually supports a claim still requires human/agent judgment.

Adjacent fields

Key voices

Latest evidence

Recent additions

All synthesized insights →

Shahaf Antwarg

How I Work With Coding Agents

A workflow synthesis from a developer running 20 parallel coding agent sessions: role shift to tech-lead, isolated worktrees per task, five-part first-message discipline, early direction-checking over end-diff review, manual commits, and a living instructions file.

dex

show-me: Visual Agent Communication & Why Software Factories Fail

Dex introduces /show-me, a skill for making coding agents communicate via compact visuals (component trees, call stacks, diagrams, pseudocode) instead of prose walls, then argues in a companion essay that 'lights-off' AI software factories fail because models can't maintain codebase quality over time, making human-in-the-loop program design and vertical slicing essential.

Adrian

Fable Planning Recipe for Autonomous Software Factories

Adrian outlines a structured planning workflow using Fable to front-load work before large-scale unattended agent execution: idea capture, adversarial questioning, research subagents, PRD drafting, policy setting, task decomposition into parallelizable 'beads', and adversarial review of tasks against the PRD.