Project Factory postmortem: 501 apps shipped, and the 6 failure modes that almost stopped it
A while back I wrote about the Project Factory — an autonomous pipeline that scaffolds, builds, verifies, and deploys small web apps with almost no human in the loop. That was the launch announcement. This is the other half of the story: the closeout.
Final verified state: 501/501 projects deployed, 0 needs-human, all thumbnails serving. The gallery is live, the stats check out, and the project earned its slot on the portfolio. Getting there meant debugging six distinct failure modes, none of which was in the original design. Every one of them taught something worth writing down.
1. The supervisor that respawned the orchestrator mid-write
The orchestrator daemon ran under a process supervisor configured to restart it on exit. Reasonable — except the orchestrator periodically rewrote a shared projects.json state file, and a restart landing mid-write meant two processes briefly believed they owned the state. The result was a classic lost-update race: progress records from one run silently overwrote another's.
Fix: single-writer discipline. Only the orchestrator process writes state; the supervisor's job is restarts, not reads. Restarts now happen only at safe checkpoints, and state writes are atomic (write temp file, rename over the original) so a crash mid-write never leaves a half-written file behind.
Lesson: supervisors restart processes; they don't serialize them. If your daemon owns a state file, the restart policy is part of the correctness design, not just ops.
2. The retired model ID that killed workers silently
Nine projects stalled with no error anywhere. The workers exited, wrote no result file, and the daemon saw… nothing. No failure, no retry trigger — just absence.
Root cause: the Tier-1 model ID in the worker config had been retired upstream. The provider returned an error the worker didn't recognize as retryable or fatal, so it exited 0 with no output. Silent success on total failure — the worst kind.
Fix: treat "no result file" as a failure state, not an unobserved one. The daemon now reconciles expected vs. actual result files every cycle, and any worker exit without output is retried and then escalated. Model IDs moved into config that's validated at startup against the provider's live model list.
Lesson: never trust exit codes from LLM workers. Verify the artifact exists, or it didn't happen.
3. Timeouts that never advanced tiers
Slow builds on free-tier models kept timing out — and then retried at the same tier, timing out again, until they were parked after 3 attempts. The tier system existed precisely for this (slow work escalates to a faster tier), but the timeout path never invoked it.
Fix: timeouts now advance the tier instead of burning a retry at the same one. Three timeouts still parks the project, but only after it's had a shot at every tier.
Lesson: every retry path needs to answer "retry differently, or retry identically?" Identical retries on a deterministic failure are just slow failure.
4. Two stranded projects the daemon never picked up
Two projects sat at attempts: 5 indefinitely. Not failed, not queued — invisible. The daemon's pickup query selected projects below the attempt cap, and these two had exactly hit a boundary the query excluded. They weren't in any bucket: too many attempts to be picked up, too few signals to be flagged.
Fix: the pickup query now uses explicit states instead of numeric boundaries, and a periodic sweeper lists anything that hasn't changed state in N cycles regardless of counters. Counters inform priority; states drive scheduling.
Lesson: boundary conditions in scheduler queries create limbo states. Enumerate states explicitly, and add a sweeper for anything that falls through.
5. The deploy death-loop that was actually RAM
Deploys kept dying and restarting in a loop. All signs pointed at the deploy script — until nothing in the script explained it. The actual cause was resource contention: the build box had 8 GB of RAM, and parallel builds plus the orchestrator plus the OS tipped it into swap, then into the OOM killer, which took out a build mid-deploy, which the supervisor restarted, which contended again. A death-loop with no bug in any single component.
Fix: cap build parallelism to fit the box, and give the deploy step a memory ceiling below the OOM threshold. Not glamorous, but the loop stopped the same day.
Lesson: when the failure is cyclic and no component is at fault, look at the substrate. Contention bugs don't live in code — they live between processes.
6. Thumbnail backfill for 23 projects
Twenty-three shipped projects had no thumbnails — not failures, just gaps from an earlier thumbnail module version. The gallery looked broken despite 100% deploy success.
Fix: a fixed thumbnail module plus a backfill pass over exactly those 23, run while the live orchestrator kept working with the old module. Backfills should be idempotent and runnable alongside the live system, not as downtime.
Lesson: "shipped" and "complete" aren't the same. Define done to include presentation artifacts, or schedule the backfill explicitly.
The numbers, carefully
Earlier snapshots showed intermediate states (478/501, 21 needs-human). Those were real at the time and are not the final state. Final verified state: 501/501 deployed, 0 needs-human, all thumbnails serving. If you're building something like this, snapshot your counters with timestamps — mixing numbers from different days is the easiest way to lie to yourself in a postmortem.
What I'd do differently
Validate external IDs at startup. Verify artifacts, not exit codes. Make retries escalate, not repeat. Enumerate scheduler states. Respect the box you're on. And define "done" to include the thumbnail.
The factory works. These six bugs are why.