Eleven days, 136 issues: what happens when agents actually run a studio's tracker
Starting 2026-07-23, every piece of work for a one-person-plus-designer studio — its Owner, its Operator, and a roster of AI agents — got routed through a single Paperclip issue tracker. The question behind it, tracked as LAB-138: can thirteen agents actually run a studio, and where does it break? By 2026-08-03, eleven days in, the company had produced 136 issues. Here's what a direct read of that data shows.
The roster said thirteen agents. The company's agent list — Tech Lead, Marketing Lead, Content Writer, Social Media Manager, Growth, Frontend Dev, Full-stack Dev, Platform Dev, Game Dev, Ops, Challenger, Reflection Coach, Summarizer — has exactly thirteen entries. Three of them, Growth, Reflection Coach and Summarizer, show a null last-heartbeat for the entire window: they never ran once, and zero issues were ever assigned to them. "Thirteen agents" on paper was about ten in practice.
Work wasn't spread evenly across the ten that did run, either. Issue counts per agent for the window: Tech Lead 35, Ops 12, Marketing Lead 12, Full-stack Dev 10, Platform Dev 9, Challenger 9, Frontend Dev 7, Game Dev 6, Content Writer 2, the rest 0. Tech Lead alone carried 35 of the 136 issues — 26% of everything that moved through the tracker. A structure meant to be distributed behaved like one coordination hub with spokes.
Of the 136 issues, 108 closed done (79%) and only 4 were cancelled (3%) — most work finished rather than got abandoned. But 75 of the 136 (55%) had no project attached at all: pure coordination, approvals, process. Only 61 (45%) were product work, spread across projects that range from a fake-commerce shopping simulator (Hayalet Sepet) to a live public-interest site (a boycott tracker) to a handful of SaaS tools still pre-launch. More than half the tracker's throughput was the studio running itself, not shipping product.
The sharpest bottleneck was a human decision, not an agent one. A status check on 2026-08-02 found six pull requests waiting on a merge call: three delivered that same day (LAB-126, LAB-127, LAB-128, including PRs on the NicheFinder and boycott-tracker projects), three waiting roughly three days since 2026-07-30 (LAB-119, LAB-118, LAB-117, again including NicheFinder and boycott-tracker work). An agent comment from 2026-07-30 on one of the waiting PRs (LAB-118) reads: "waiting on your merge decision on GitHub... per the 'merge kararı Overseer'da' rule" — roughly, "the merge call belongs to the Overseer." The agent had finished the work and had nothing left to do but wait. That pile-up is the direct reason the human merge gate was removed the same day: agents now merge their own PRs once checks are green.
Removing a rule and executing it turned out to be two different events. As of this writing, after the gate was gone, four PRs (LAB-119, LAB-126, LAB-132, LAB-118) were still sitting open and unmerged, still marked in_review — a live check on 2026-08-03 confirmed it. Nothing carried them through automatically; the queue needed someone to wake the responsible agents and point them at the now-unblocked work. A policy change in a document is not the same event as the queue actually clearing — arguably the sharpest single finding of these eleven days.
A second, different kind of gate hasn't moved at all. The boycott-tracker site's launch has been stuck on a single step since 2026-07-27 19:14: paste three Auth0 environment variables into the hosting panel, redeploy, run a login smoke test — plausibly ten minutes of work. The system's own response to the wait has been a daily auto-generated "review productivity" child issue, opened and closed seven times between 2026-07-27 and 2026-08-02, each one re-confirming the same three facts: the site returns 200, TLS is valid, the blocker hasn't moved. Eleven such productivity-nudge issues were generated across the window overall, seven of them for this one task. The system's only self-generated response to a stuck human gate is a polite reminder loop — there's no escalation built in.
The size of that decision queue is easy to undercount if you only look at direct assignment. Six issues were assigned straight to the Owner over the eleven days, five to the Operator. But the real decision load runs higher than that count suggests: the boycott-tracker launch alone needed five separate human approval points, not one. A snapshot taken while writing this showed twelve issues sitting in_review or blocked at the same moment.
On the studio's anti-goals: a full-text pass over all 136 issue titles and descriptions found zero agent-initiated issues mentioning revenue, monetization, pricing, or advertising. The one monetization-related task in the data belongs to a separate, paused project and was opened directly by a human — not an example of an agent proposing scope on its own. Worth being honest about the limit here: this shows no example surfaced in eleven days, not that the rule was tested and held. There's no way to tell, from this window, whether the opportunity never came up or the constraint is doing real work.
One more small thing. At least three NicheFinder issues in quick succession (LAB-119, LAB-126, LAB-132) had agents discover mid-task that their written spec no longer matched the code: one referenced a Supabase RPC after the project had already moved to Prisma and Postgres, another described a Favorites API a previous pull request had already shipped. Parallel agents move fast enough that written task descriptions fall behind the codebase within days. Not a failure, just friction a faster loop produces on its own.