Syllogistic Software Inc.

Can AI write our software without shipping broken work?

Yes, if the coding agents work inside a process that catches their mistakes before a customer does. On its own, an agent will sometimes write code that looks right and is wrong. Put it through a fixed lifecycle of specification, plan, tests, an independent review and a human approval before merge, and the broken work is sent back before it ships rather than after.

Evidence

I build my own software this way with WASBuilder, a board where coding agents specify, plan, implement, review and stage each ticket, and a human answers their questions and approves the merge. Every ticket keeps its specification, plan, acceptance criteria and test plan in plain text, and every agent run records its model, time and tokens.

A WASBuilder ticket page showing the specification, plan, acceptance criteria, and test plan
A ticket page with its specification, plan, acceptance criteria and test plan. This screenshot is from a demo project with example data.
The review report, notes, and agent runs on a ticket, each run with its model, duration, and tokens
The review report and agent runs on a ticket, each run with its model, duration and tokens. This screenshot is from a demo project with example data.

The numbers below are real. They come from the production board and cover two of my projects, WASBuilder itself and TaskTree, from 2026-08-20 to 2026-10-08:

  • 220,486 lines of code in two projects built with WASBuilder, current codebase size counting all files except shared libraries
  • 599 of 614 tickets created were shipped to production
  • 17% were sent back by Final Review, the independent agent review, before they could ship
  • 18% needed any rework at all, counting review, a human sending them back and a failed staging deploy
  • 254 lines changed in the median ticket

Read the full story in the WASBuilder case study.

What it takes

A codebase the agents can build and test with one command, because the tests are what tell an agent its work is broken. Without them, review is guesswork.

Someone on your side who knows what the software should do. The agents stop and ask when a ticket is unclear, and they need a person to answer and to approve each merge. That is not a full-time job, but it cannot be skipped.

Small tickets. The process works best on changes with a clear outcome. Large, vague requests get split up first.

The honest limits: it is not perfect. Review catches a lot, and a human still has to read what is about to ship. A part of your system with no tests, or with rules that live only in someone's head, is where broken work is most likely to slip through.

Next step

Tell me about your codebase and team, and I will tell you plainly whether coding agents can work on it safely and what would need to change first.

Sign in

or

or
Sign up

or
Account
Change email address:
Enter current password:
Change password: (blank to leave unchanged)