AI Development

The talk: 2026, the year of the software factory

16 August 2026

What I argued at DDD South West in May — the vibe coding hangover, the return of the software factory, and the evidence slide with the failure left in — plus what three more months of running the thing have taught me.

On 16 May I gave a talk at DDD South West, at the Engine Shed in Bristol, titled “2026: The Year of the Software Factory”. Here’s the write-up — it has taken me rather longer than planned.

I had originally meant to talk in fairly abstract terms: vibe coding, agent swarms and software factories as three approaches, with some signposting about where things looked to be heading. But the more of the factory I actually built, the more the scope shifted, and it ended up being a talk about what I had done and what I had learned doing it, with an attempt to point out the things I thought mattered.

The argument

Karpathy coined the term “vibe coding” in February 2025, and by early 2026 he had moved on from it himself in favour of “agentic engineering”. That shift, from vibes towards engineering, was essentially what the talk was about.

The industry figures I opened with suggest AI-assisted projects tend to ship faster but with more issues and more technical debt. My own view is that tools optimised for the first 80% of a project can leave teams stranded in the remaining 20%, which is usually where the difficult work is.

Software factories themselves are not a new idea. Bemer at GE and McIlroy at AT&T were arguing for them in the late 1960s. What has changed is that a repeatable, structured process with verifiable outputs is now achievable for one person on one machine, because the workers can be agents.

The bulk of the talk was a walkthrough of mine: an intent router, an orchestrator, an architect agent, a product owner agent, and build agents working through an adversarial Player–Coach loop, all running on a DGX Spark under my desk. It represented about nine months of evenings and weekends at that point, and I presented it as a case study rather than as a product launch.

The evidence slide

The slide I was most interested in showing covered a comparison we had run. A fine-tuned Gemma 4 26B, trained on 2,200 examples harvested from nineteen architecture books by the factory’s own dataset pipeline, reviewed the same real project architecture documents as GPT-5.5. It matched the larger model on three of the four sessions, at $0.00 a session against $0.15–0.30, and was faster too once the model was warm.

Session 1 failed, because of a schema gap in the training data, and I left that on the slide. Removing the failures would have turned an evidence slide into a marketing one.

The economics part of the argument was changing while I was preparing it. In the three weeks before the talk, one vendor announced it was pulling programmatic access from its subscription plans and another moved to usage-based billing, so I was still updating slide 6 three days before speaking.

On the day

I made it hard work for myself. I had been building so much that I didn’t start on the slides until the week of the talk, and I was entering the Kaggle hackathon a couple of days afterwards, with a Reachy Mini robot that had arrived a fortnight earlier and still needed integrating. It was manic.

The build had gone better than I expected, which was the problem: more of it came together, and more easily, than I had planned for, so I had far too much content for a thirty-minute slot. An hour would have been about right. I had practised the talk roughly once, then made changes to it partway through, so I was dead nervous by the time I stood up.

I had also recorded a demo at home, in case there wasn’t time to run one live — which there wasn’t. It shows a change proposal going into OpenWebUI, routed by Jarvis to the architect agent to validate, and a build then queued in the forge; followed by a study session on An Inspector Calls with the Reachy Mini.

What the room made of it

DDD South West collect feedback on the day and pass it on afterwards. Fifteen people responded, averaging 4.5 out of 5 for knowledge and 4.2 for skills, which was kinder than I felt at the time.

The comments were more useful than the scores. One person wrote “Needs to be a full hour talk!”, which is exactly what I had concluded walking off. Several said it was pitched above them — “possibly a bit over my head”, “I did get a bit lost with some of the concepts” — and one noted the session description had made it sound lighter than it turned out to be. Someone couldn’t read the slides: “a lot of information on the slides was very small”. And the one I found most useful, because I hadn’t seen it coming: “Was missing a bit of an intro — didn’t really get what he was aiming to achieve with what he was talking about.”

The comment that stung slightly was “A demo would have helped reinforce the slides.” I had one. I just hadn’t left myself the time to show it.

Three months on

The talk describes where things stood in May, and the factory has moved on since. Most of the work since has gone into verification: the BDD layer meant to prove each feature worked had become a nightmare of broken glue and plumbing, and replacing it mattered a great deal more than adding any further agents. The first feature to run the whole chain — from a written sentence through to merged code, with a single human pause in the middle — was on 15 August.

I closed the talk with the line “The AI is the engine. You’re still the engineer”, and three months of running the system has, if anything, reinforced that view. There is a fuller account of where the factory has ended up in the software factory case study, including the deck.

Newsletter

Get new posts by email

Occasional write-ups on agent systems, evals and local models — including what broke. No pitch, unsubscribe in one click.