Skip to content
LiveGraph

Our engineering team is a LiveGraph graph. Here is what broke.

October 4, 2026

Every pitch deck says “we use our own product.” Ours is more literal than most: the repository this site ships from is improved, reviewed, and partially maintained by a LiveGraph graph running on the hosted LiveGraph instance. Not a demo graph — the same one whose runs you can watch on the canvas, whose parked approvals sit in our /approvals queue, and whose branches land as real pull requests against the code you are reading right now.

This post is the architecture and the failure log. The failures are the useful part — running agents unsupervised for months teaches you things a weekend eval never will.

The shape of the loop

A scheduled trigger (loop-tick, every six hours, pinned) fires the Lead — the graph's entry node — with the loop brief as its input. The Lead hands the round to the Loop Engineer, which works a queue file (QUEUE.md) in the repo: pick the top open item, implement it, write a report. An overlap guard serializes rounds — a tick that fires while a run is still active just skips, at zero cost.

The engineer has no checkout and no filesystem. Everything goes through the GitHub MCP server: get_file_contents and search_code to read the code, the queue, and the repo's own docs; create_branch plus push_files/create_or_update_file to land the smallest possible diff on a livegraph/* branch. Working on live remote state turned out to be a feature: a volume-mounted clone can go stale; the remote can't. The trade-off we accepted — writes are whole-file, and the branch + CI + PR is the review surface instead of a local worktree.

There is deliberately no edge from Engineer to Reviewer. A review that fires when the push lands would pre-date CI. The only path to review is a webhook.

CI routes its own verdict — before any model runs

Every push to a livegraph/** branch runs a deliberately cheap CI workflow: schema drift check, unit tests, both typechecks, web lint. No secrets, no billed end-to-end suite — this workflow exists to run on branches an agent can create, and the billed suite stays the human gate on main. When the run completes, GitHub fires a workflow_run webhook back into LiveGraph, and the webhook trigger's CI route decides the destination in the route, before any hop:

  • Failure on a loop branch → a new run starts at the Loop Engineer — a fix round on the same branch. The routed input carries the failing run's id and URL, and the engineer reads the real job logs over GitHub's API before touching anything. One informed fix attempt per input. A per-branch cap on fix rounds keeps an unfixable branch from looping forever.
  • Success on a loop branch → a new run starts at the Release Reviewer.
  • Anything else — another branch, another event — is a 202 skip. No run, no billed hop.

The reviewer reads the diff, not the repo

The Release Reviewer's instruction is one line of prompt doing real work: the diff already carries the verdict. It calls pull_request_read (get_diff — one call, the whole unified diff) when a PR exists, else list_commits + get_commit for the same per-file hunks. get_file_contents is reserved for the file an item's acceptance line names, or a hunk too small to judge — a rename, moved code, a patch GitHub elided. The budget is written into the prompt: the PR check, the diff, and at most ~3 targeted file reads.

An APPROVED verdict routes to a thin dispatcher node that hands off to a separate Publisher graph — deliberately separate, so a parked approval lives in /approvals instead of stalling the loop's overlap guard for up to 48 hours. A human approves (or edits, or rejects with a reason), the publisher opens the PR, and GitHub auto-merge lands it when the required checks go green. The publisher has no merge tool — it cannot force a red merge.

What actually broke

Three honest failures, each now structural:

  • Whole-file reads burned $2.22 on one run. An early engineer on a pricier model re-read entire files to orient — and one get_file_contents call on a large source file resends the whole file at input price on every subsequent step of the tool loop. One read once ate a 200k context window outright. The fixes landed in three places: a read-range tool so slices by line number are the taught default, the diff-first reviewer budget above, and a per-run dollar ceiling stamped on the graphs so a runaway tick has a hard wall.
  • Model-triaged routing dropped real work. Before the CI route existed, the Lead read each webhook event's JSON and decided where it went. It answered DONE to a CI success on a loop branch — stranding the review entirely — and every CI verdict on a non-loop branch was a billed no-op hop. Routing a verdict that carries its answer in one field is a code problem, so it moved into code: the route itself maps conclusion + branch to a destination. Auto-routing had its own subtler bug in the same neighborhood — a delegation message that mentioned another node's name mid-sentence (“Release Reviewer — a Platform Engineer run committed…”) let that node's name outscore the intended target. The resolver now treats a specialist's name at the start of a reply as decisive and ignores mid-text mentions.
  • Invalid request bodies returned bare 500s. Our API parses bodies directly in each route rather than through schema-validated handlers, so a thrown validation error was an uncaught exception — and the global error handler that mapped it to a 400 was registered after the route plugins, which under Fastify's encapsulation model means the routes never saw it. Every malformed body in the app became a bare 500 until an e2e test expecting a 400 caught it. The fix was registration order — and the lesson generalizes: the checks you bolt on last are the ones every caller silently misses.

What running it unsupervised actually taught us

The pattern across all three: move verdicts into code, and budget the model's reach. The CI verdict routes in the route, not in a prompt. The publish tier is classified by a tool diffing the actual changed-file list, not by the reviewer's label — and the ungated fast lane re-verifies that in code before it publishes. The queue file's edits are checked by a deterministic diff tool, not a careful reviewer. What's left for the model is exactly what models are good at: reading code, writing diffs, exercising judgment on a change — with every irreversible step (the PR, the merge, anything outbound) behind a human or a check that doesn't speak model.

The loop is still running. Its reports land in docs/loop-reports/ on its own branches, its fix rounds self-heal failed pushes, and roughly everything it's shipped went through the same approval queue a customer's runs park in. If you want the same shape on your own repo — a queue file, an engineer, CI routing, a reviewer, a human gate — it's the Improvement Loop template: point it at a repo, and watch the first tick on the canvas.

See it running

The live demo needs no account and no API key — a scripted model drives the real engine while you reroute a run in flight. When you want your own graphs, sign up and the included model runs them; bring a key when you want a different provider.

Keep reading

  • We gave LiveGraph a paragraph. It built citepath.ai — and it hasn't stopped. — citepath.ai is a live GEO SaaS built entirely by a LiveGraph bootstrap run and its improvement loop: provisioning, self-hosted CI, thirty-three merged PRs, every failure — and what each failure became.
  • A loop wrote a complete SaaS for $23.85 — then handed over the keys. — findefend.com went from a paragraph-long brief to a live, self-improving product: one bootstrap run, fourteen reviewed PRs, real incidents that became platform fixes — and a credential handoff where the model never saw the password.
  • Share a run people can watch, not a log they can't — POST /runs/:id/share mints a stateless token that renders a read-only, live-updating canvas — topology and hop status, never hop content. Embed it in docs, status pages, or a client deliverable.
  • LangGraph edits a run's state. LiveGraph edits a run's route. — LangGraph compiles your topology into code; LiveGraph keeps it in data you can mutate between hops. Two different primitives — an honest map of where each one fits.
  • We pointed LiveGraph at an empty domain. It shipped a product. — modelright.dev went from a domain name to a live, self-improving app through one bootstrap run and a 10-minute improvement loop — approvals, conflicts, and all.
  • Steer a running agent workflow — LiveGraph re-reads the graph after every hop, so dragging an edge changes where this run goes next.
  • Why you can't reroute a running n8n workflow (and why you can in LiveGraph) — Most engines compile a run before executing it, so mid-run edits only affect the next run. A per-hop engine makes rerouting free — here's the architecture.
  • Giving agents write access to real code without losing sleep — Worktree isolation, operator allowlists, encrypted secrets, and pushes only a human's literal keystrokes can trigger — the safety model, decision by decision.

LiveGraph

Model-agnostic agent orchestration on a live canvas. Hosted at livegraph.ai.

Product

  • Live demo
  • Pricing
  • Security
  • Automations
  • Compare
  • Migrate from Flowise
  • Migrate from Agent Builder
  • Sign up

Resources

  • Blog
  • Docs
  • Changelog
  • Model Radar
  • Status
  • RSS

Agents

  • Agent guide
  • llms.txt
  • llms-full.txt
  • MCP integration

Support

  • [email protected]

Legal

  • Privacy
  • Terms
© 2026 LiveGraphSite by Canweb Ltd.