DesignX · AI-native design orchestrator
Design as an agentic pipeline, not a prompt.
An orchestrator that takes whatever project knowledge you have and returns a design system of record: every screen cited, every value a resolved token, every phase held at a gate until a human clears it.
My role
Founding member. I shaped the pipeline and its gates, wrote the 11 agents and 97 skills that run it, set the quality bar it grades against, and designed and built every screen of the product.
Ina regulated fintech brief
Outa design system of record
- P0Onboard
- P1Discover
- P2Ideate
- P3Design
- P4Deliver
- P5Measure
The contract
- cited
- tokenised
- graded
- buildable
The input varies. The contract does not.
The gap
AI made design faster, and impossible to defend.
By 2026 generative tools could produce a screen in seconds. The market split into two camps, and neither one produced the thing a senior team actually needs.
One-shot generators
Copilots
A system of record
Undefendable work
There was a third option, and it did not look like a chat box at all.
The bet
What if design ran like a production line?
Not a prompt you fire and hope. A governed pipeline: distinct phases, specialised agents, hard approval gates, and an evidence contract where nothing ships unsourced.
Before · one leap, no lineage
“make me a grocery app, make it nice”
After · every step gated, cited, reviewable
Discover → Ideate → Design → Deliver → Measure. Each boundary is a gate a human has to clear.
My role
I owned the line, from strategy to the last pixel.
I was a founding member, there from inception. I did not join a team building this; I helped decide what it was.
I shaped the strategy and the pipeline itself, wrote the agent and skill definitions that make it run, set the standard its output is graded against, and built the eval harness that decides whether work is good enough to clear a gate. The calm surface hiding all of that machinery is mine too, end to end.
What I owned
Depth
- Product strategy and pipeline shape
- The six phases, the nine gates, and the reframe from prompt to production line.
- Led
- Agent and skill definitions
- 11 agents and 97 skills across 13 categories, plus the protocol every agent inherits.
- Led
- The quality bar
- A 570-line design authority that grades every visual output, and the eval harness that enforces it.
- Led
- Orchestration and control model
- Run modes, autonomy levels, model routing and context budgets.
- Co-led
- The entire product UI
- Every screen in this case study. Design and build, start to finish.
- Sole
Sole means exactly that: no other designer touched the product surface.
The pipeline
Six phases. Nine gates. Nothing skips.
11 agents and 97 skills sit behind six phases. A phase cannot advance until its gate clears, and a human decides when it does. There is no auto-advance anywhere in the framework.
Tested against 11 project shapes
Brownfield codegen, Figma import, consumer web, enterprise B2B, greenfield mobile, hybrid IoT, regulated fintech, design-system refresh, research sprint. The same gates apply to all of them.
Three calm questions and a knowledge-base upload. Name it, pick a run mode, drop in what you already know. Everything the pipeline needs, before any generation happens.
Agents read the uploaded corpus and produce insights with evidence IDs and confidence levels attached. Nothing is asserted without a citation back to a source in the knowledge base.
Concepts are generated against the insights, then scored against a requirements checklist. On the GreenBasket run: 14 of 14 requirements met, zero blockers, before the gate opened.
A design system of record is produced first: tokens, components, states. Screens are then specified against it, and a Design Director pass reviews visual quality before anything is called done.
The Builder turns the design system and screen specs into real, running code, opened in a live preview beside the conversation. Spec and product stay in the same window.
A designer-QA rubric scores the output across eight dimensions. The GreenBasket run scored 4.1 out of 5 and surfaced 16 gaps, all of which were resolved to production specs.
Progressive disclosure
Hiding a studio behind three questions.
The hardest design problem was not the machinery. It was making the machinery feel like nothing. Every one of the 97 skills stays invisible until the moment it is needed.
- 01
Name it in a sentence
What are we building? A name and a description. That is the entire first screen.
- 02
Pick a run mode
Guided, Balanced, Autonomous or Express. How much of the wheel you want to hold.
- 03
Drop in what you know
Briefs, transcripts, surveys, brand files. Up to 50 MB. The pipeline reads all of it.

Then, and only then

Control
Autonomy is a dial, not a switch.
Trusting a pipeline is not all-or-nothing. Four run modes move the same pipeline along a single axis: how often it stops and asks you. The gates never disappear; they only change who opens them.
Run mode
Gate behaviour
Phase boundaries pause
When you reach for it: The default.
Filled = pauses for you · Hollow = clears itself
Control is only half of trust. The other half is proof.
The eval
A rubric that tells you where it is weak.
Generation is easy to fake and hard to trust. So the pipeline scores its own output against a designer-QA rubric across eight dimensions, and reports the low scores as loudly as the high ones. The GreenBasket run came back at 0.0/5 with 16 named gaps, every one resolved to a production spec before delivery.
Dimension
Score / 5
The lowest score is the point
Interaction detail scored 3. That is the number I care about most.
A tool that grades its own work at five out of five is not an eval, it is marketing. The rubric exists to produce a worklist, and a 3 is where the next run starts. Each gap arrives as a named, specified fix: not “improve the interactions”, but the component, the state, and the intended behaviour.
16 gaps found · 16 resolved to production specs · 0 blockers at the concept gate
The record
The deliverable is not the screen. It is the system that produces screens.
A one-shot tool hands you a picture. DesignX hands you the argument for every pixel in it, and the system that will produce the next hundred consistently.
Evidence
Insights
Design system
Production specs


All of which is an argument. Here is what came out the other end.
The output
From brief to a working app.
GreenBasket was a throwaway test project, a fictional London grocery marketplace, run through the pipeline to see what came out the other end. Not a mockup: a working front end, built from the generated design system, with real copy and real product logic. These are its screens, untouched.

Three decisions worth pointing at
Every claim has a “Learn more”
Four trust claims sit above the fold, and not one of them is a bare assertion. Each opens its own methodology. “No dark patterns. Ever.” is an unusual thing for a generated store to promise, and it traces directly to the INS-08 implication.

Substitution is a preference, not a surprise
Every basket line carries its own “if unavailable” choice: substitute, contact me first, or remove. Fresh-grocery delivery lives or dies on this interaction, and the pipeline surfaced it per item rather than burying one global setting in an account page.


Allergens on the card, not the detail page
Products are checked against the 14 major EU allergens, and the warning renders on the card itself. The Impact Score badge sits in the same corner on every tile: the kind of consistency that comes from specifying a component before drawing screens.


Proof
Proof it wasn’t a prototype.
We shipped DesignX to npm to test one thing: would anyone want design run as a governed pipeline? They did. Then we took it private once the idea was proven.
Public run
0
downloads in the first year
- Releasesv1.0.0 → v5.5.3
- 52
- Major versionsin ~3 weeks
- 4
- Peak daylaunch
- 3,237
One proof run, end to end
- Knowledge base in
- 16 KB
- Artefacts out
- 134
- Concept gate0 blockers
- 14 / 14
- Designer-QA score16 gaps, all resolved
- 4.1 / 5
- Total model spend
- $5.78
The lasting idea isn’t the agents. It’s that when design becomes a pipeline, the deliverable stops being a screen.
It becomes the system that produces screens, one where every decision can be traced, defended, and re-run. That is the difference between a tool that makes design faster and one a senior team can actually stand behind in a review.



