Your coding agent ships code.
This is what checks it.
You write a spec. Your Claude Code builds against it, a second model from another family reads the same diff, and where they disagree you see the disagreement and decide. Three decisions stay yours — what gets built, how, and whether it ships — and spending caps refuse rather than warn. What lands is a repository you own and a URL that works. Requires Claude Code, or Codex for the skills and MCP server. Median $171 in tokens across seven products, measured 2026-07.
— downloads on npm15 industries · 60 products · 6 reusable pipelines
The 60 products collapse into 6 reusable build pipelines (CRUD vertical-SaaS, booking, CRM, dashboard, marketplace, content) — each ships through three default stops — one line in PROJECT.md takes it to one — then automated to a live URL. See how it works ↗
Next.js, Postgres, shadcn, Stripe, Tailwind — GreatCTO builds on the tools your team already trusts, and ships a repo you fully own.
Building a product is now three approvals.
The next wave isn't a faster code editor — it's a builder that takes a product from idea to shipped software. You describe what you want; specialist agents handle architecture, the data model, the backend, the frontend, the tests and the deploy. The three things you do are approve what gets built, approve the design and approve the deploy. Everything between runs unattended, to a repo you own and a live URL. The first of the three is new in v3: the pipeline used to stop on how to build and whether to release, and never on what to build — the decision that is wrong for six stages before anyone finds out. It costs one pause per product, not per feature. And every model improvement makes the build faster and cheaper — the economics only bend your way.
A pipeline of specialist agents — with one gate that's yours.
Spec → build → test → deploy.
A chain of specialist agents runs the real steps of shipping software — architect, design, build, QA, deploy — not one chat window. You watch the whole pipeline on one board.
You approve what, how, and ship.
The first checkpoint: the architecture and plan. Approve it and scaffold, backend, frontend, integrations, tests and deploy all run automatically — no per-line babysitting.
A repo and a live URL.
Not a prototype in a sandbox — a working product: a git repo you own, a deployed app on Vercel or Cloudflare, and the tests that keep it honest.
Generated tests, green before ship.
Every build writes its own tests and must pass CI. You sign the direction; the test suite catches the regressions — so automation stays safe without a human reading every diff.
Next.js, Postgres, shadcn, Stripe.
Built on the tools your team already trusts, with a React Native option for mobile. No bespoke framework to learn — it's a codebase any developer can pick up.
One template ships many products.
The 60 products collapse into 6 build archetypes — CRUD, booking, CRM, dashboard, marketplace, content. See the pipelines ↗ Self-hosted, MIT, your code stays on your machine.
What do you want to build?
Pick a product — see the pipeline, the agents at each stage, and the one CTO gate.
Describe it. Approve the spec. It ships.
Say what you want to build.
A dispatch app, a booking portal, a CRM, a dashboard — name the product and the industry. The architect and design-advisor draft the spec, data model and screens.
The spec is where you sign off.
You review the architecture and plan, and sign off. That is the second of three checkpoints by default — the brief came before it, the deploy comes after. At the lightest setting this one is a screen you read, not a form you sign.
Build → test → deploy, automated.
Scaffold, backend, frontend, integrations, generated tests and deploy run end to end — to a repo you own and a live URL.
Automated to the deploy.
The rails keep it honest.
Letting a build run unattended is only safe with rails. Every pipeline ships with three: human checkpoints you choose the depth of — three by default, one if you want — generated tests that must pass CI before anything ships, and a repo you fully own and can read, diff, and roll back.
You set the depth.
The default stops three times and nowhere else: what gets built, how, and the deploy. One line takes it to one — you approve the deploy, and what gets built arrives as a screen you read instead of a form you sign. Silence is consent, and the screen says so. Nothing irreversible runs unattended at any setting.
Generated tests, green before ship.
Every build writes its own tests and must pass CI. The suite — not a human reading every diff — catches regressions, so the automation stays safe at speed.
Readable, diffable, reversible.
The output is a normal git repo on a modern stack — not a black box. Review the diff, run it locally, fork it, or roll it back. Self-hosted, MIT, your code stays on your machine.
A CTO dashboard — that runs itself.
great-cto board opens a live board on your own machine — the pipeline, cost, and project memory in one place. No account, no SaaS, no telemetry by default. And you never set it up: it fills in as you work.
No /audit, no /save.
The first session maps your codebase automatically; every agent run records a verdict that feeds the metrics; every session auto-saves a log and extracts lessons on exit. You just work — the memory, metrics, and logs fill in.
Runs on localhost, yours.
No sign-up, no seats, no cloud dashboard. It runs on your machine against your repo, offline-capable, and your code stays on it. Open source, MIT.
Pipeline, cost, memory.
The live pipeline with its risk-tier gate, per-agent cost, 30-day LLM spend with its provenance — measured, estimated, or unmeasured, and browsable project memory — PROJECT.md, archetypes, lessons. Documents are searchable inside, not only by name.
A second model reads the work.
The pipeline used to hand one stage's output to the next on the strength of a line the stage wrote about itself. Now a different model checks it — cheapest question first: do the named files exist, do the frozen criteria pass when run, and only then is a model asked. Three answers, never two: verified, send-back, or unverifiable — which is not a pass, or the cheapest way to pass would be to claim nothing.
Quality is its own record.
A verdict says what a run did. A score says how well — separately, so a later check can disagree with an earlier one without rewriting it, and so every assessment names who made it. A run nobody assessed counts as nothing, never as zero: a pass rate divides by what was actually looked at.
A budget that refuses.
Cap what any agent may spend, from the board. Past the cap the pipeline declines to dispatch it and names the number — it does not fail quietly. And an estimate never refuses: while no run has recorded a real cost the cap reads unmeasured and holds nothing, because a limit firing on a number nobody measured is worse than no limit.
Even doing nothing is recorded.
Every dispatcher decision is written down, including the ones where it decided not to act, and why. "Nothing should happen" and "nothing could happen" look identical from outside — that gap is where pipeline defects live, and it is closed.
Tokens per build.
Not seats. Not agencies.
A build costs the tokens it uses. What that is for your product, the board measures; we do not estimate it for you.
Pay for the work, not seats.
Every model improvement makes the same product faster and cheaper to build. The economics only bend your way.
MIT. Self-hosted. Yours.
You run it; your data and your repo never leave your machine. No per-seat tax, no SaaS lock-in.
You pay your tokens — we don't bill you.
No GreatCTO invoice. Bring your own Anthropic / OpenAI key.
One feature, end to end:
1h 26m and $3.40 in LLM cost.
A real run, fully public: spec → build → review → tests → merged PR. Every stage timestamped, every artifact links to a real GitHub PR — no screenshots, no marketing math.
median $171 · 70/100
The open benchmark built 7 products end to end: median $171 in tokens, median quality 70/100 (range 58–86). Reproduce it with scripts/bench-run.sh.
1h 26m · $3.40 LLM
Architect → plan → implementation → review → tests → merged PR. Wall-clock from prompt to ship, with one human signing the spec.
Timestamps, PRs, costs.
The full stage-by-stage timeline with public GitHub links. Walk the run on /proof →
Frequently asked.
What is GreatCTO?
What can it build?
How much is automated?
What does it cost?
Where does my data go?
How do I start?
Describe the product.
Ship the software.
Open source · MIT · self-hosted · your code stays on your machine