The Farm

A fleet of agents that ships software around the clock.

The Farm is the production line behind every Creative Code deployment. Agents take scoped work from a queue, build it in isolation, prove it with tests and evals, and hand it to an engineer for the final call.

farm / run-log live
09:14:00 queue pulled task #4812 scope: inventory-sync retries
09:14:23 agent-07 sandbox ready worktree feat/4812
09:14:46 agent-07 context loaded 41 files, 3 ADRs
09:15:09 agent-07 implementing claude / planner

Illustrative run log. Task names are anonymized.

What is literally running.

No mystery. A queue of scoped tasks, a pool of isolated agent workers, a control plane that holds credentials, budgets, retries and logs, and a gate that every change must pass.

  • Up to 48concurrent agent sessions coordinated across projects
  • 4model families: Claude, GPT, Gemini and Grok, routed by task
  • 70+MCP tools our agents and our customers' agents call in production
  • 1shared knowledge base, so no agent re-learns a decision we already made

Nine steps. Every change.

FORGE, our orchestration pipeline, is deterministic on purpose. Creativity belongs in the code, not in the process that checks it.

  1. 01

    Load context

    Repository, architecture notes, the brief and prior decisions from our knowledge base.

  2. 02

    Sandbox

    A fresh git worktree and container per task. Agents never touch a shared branch.

  3. 03

    Implement

    The planner model writes the change; tools are scoped to the task.

  4. 04

    Validate

    Lint, types, unit and integration tests, plus the workflow's eval set.

  5. 05

    Cross-check

    A model from a different family reviews the change. Correlated blind spots are the enemy.

  6. 06

    Fix CI

    Failures loop back to the agent with the logs, within a retry and cost budget.

  7. 07

    Security

    Secrets, dependency and static-analysis scans on every diff.

  8. 08

    Pull request

    A reviewable PR with the reasoning, test evidence and eval scores attached.

  9. 09

    Engineer decides

    An FDE approves anything consequential. Low-risk classes can auto-merge.

Agent roles

A team, not a chatbot.

Each agent has one job and the tools for that job only. Hard reasoning goes to the strongest model; volume goes to the fastest.

Planner

Breaks a brief into scoped, testable tasks and writes the acceptance checks.

Builder

Implements in the sandbox. The workhorse of the fleet.

Reviewer

A different model family reads every diff before a human does.

Tester

Writes the failing test first, then guards against regressions.

Researcher

Reads docs, specs and the web so builders do not guess at APIs.

Documentarian

Writes down every decision so the next run starts smarter.

Guardrails first

Protected files, scoped credentials, dangerous-command detection and a hard cost budget on every run.

Many models, one standard

Work is checked by a different model family than the one that wrote it. We benchmark every workflow and use the smallest model that passes.

Memory that compounds

Every decision, fix and incident is written to a shared knowledge base the whole fleet reads before it starts.

Put the Farm on your backlog