DesignX · AI-native design orchestrator

Design as an agentic pipeline, not a prompt.

An orchestrator that takes whatever project knowledge you have and returns a design system of record: every screen cited, every value a resolved token, every phase held at a gate until a human clears it.

My role

Founding member. I shaped the pipeline and its gates, wrote the 11 agents and 97 skills that run it, set the quality bar it grades against, and designed and built every screen of the product.

Ina regulated fintech brief

Outa design system of record

  1. P0Onboard
  2. P1Discover
  3. P2Ideate
  4. P3Design
  5. P4Deliver
  6. P5Measure

The contract

  • cited
  • tokenised
  • graded
  • buildable

The input varies. The contract does not.

01

The gap

AI made design faster, and impossible to defend.

By 2026 generative tools could produce a screen in seconds. The market split into two camps, and neither one produced the thing a senior team actually needs.

Camp one

One-shot generators

v0, Lovable, Figma AI. A screen in seconds, with no research lineage and no rationale, and they will invent facts to fill the gaps.
Camp two

Copilots

Genuinely useful for one artefact at a time. But what to make, in what order, gated by what: all of that still lives in your head.
Neither

A system of record

Nothing produced a durable chain: the artefact, the insight beneath it, and the evidence beneath that. Output without lineage.
The cost

Undefendable work

A team cannot ship what it cannot defend in a review. Speed that produces unciteable work is not speed.

There was a third option, and it did not look like a chat box at all.

02

The bet

What if design ran like a production line?

Not a prompt you fire and hope. A governed pipeline: distinct phases, specialised agents, hard approval gates, and an evidence contract where nothing ships unsourced.

Before · one leap, no lineage

“make me a grocery app, make it nice”

After · every step gated, cited, reviewable

Discover → Ideate → Design → Deliver → Measure. Each boundary is a gate a human has to clear.

03

My role

I owned the line, from strategy to the last pixel.

I was a founding member, there from inception. I did not join a team building this; I helped decide what it was.

I shaped the strategy and the pipeline itself, wrote the agent and skill definitions that make it run, set the standard its output is graded against, and built the eval harness that decides whether work is good enough to clear a gate. The calm surface hiding all of that machinery is mine too, end to end.

What I owned

Depth

Product strategy and pipeline shape
The six phases, the nine gates, and the reframe from prompt to production line.
Led
Agent and skill definitions
11 agents and 97 skills across 13 categories, plus the protocol every agent inherits.
Led
The quality bar
A 570-line design authority that grades every visual output, and the eval harness that enforces it.
Led
Orchestration and control model
Run modes, autonomy levels, model routing and context budgets.
Co-led
The entire product UI
Every screen in this case study. Design and build, start to finish.
Sole

Sole means exactly that: no other designer touched the product surface.

04

The pipeline

Six phases. Nine gates. Nothing skips.

11 agents and 97 skills sit behind six phases. A phase cannot advance until its gate clears, and a human decides when it does. There is no auto-advance anywhere in the framework.

Tested against 11 project shapes

Brownfield codegen, Figma import, consumer web, enterprise B2B, greenfield mobile, hybrid IoT, regulated fintech, design-system refresh, research sprint. The same gates apply to all of them.

  1. Three calm questions and a knowledge-base upload. Name it, pick a run mode, drop in what you already know. Everything the pipeline needs, before any generation happens.

  2. Agents read the uploaded corpus and produce insights with evidence IDs and confidence levels attached. Nothing is asserted without a citation back to a source in the knowledge base.

  3. Concepts are generated against the insights, then scored against a requirements checklist. On the GreenBasket run: 14 of 14 requirements met, zero blockers, before the gate opened.

  4. A design system of record is produced first: tokens, components, states. Screens are then specified against it, and a Design Director pass reviews visual quality before anything is called done.

  5. The Builder turns the design system and screen specs into real, running code, opened in a live preview beside the conversation. Spec and product stay in the same window.

  6. A designer-QA rubric scores the output across eight dimensions. The GreenBasket run scored 4.1 out of 5 and surfaced 16 gaps, all of which were resolved to production specs.

05

Progressive disclosure

Hiding a studio behind three questions.

The hardest design problem was not the machinery. It was making the machinery feel like nothing. Every one of the 97 skills stays invisible until the moment it is needed.

  1. 01

    Name it in a sentence

    What are we building? A name and a description. That is the entire first screen.

  2. 02

    Pick a run mode

    Guided, Balanced, Autonomous or Express. How much of the wheel you want to hold.

  3. 03

    Drop in what you know

    Briefs, transcripts, surveys, brand files. Up to 50 MB. The pipeline reads all of it.

DesignX new-project screen: the assistant asks what is being built, with a name field reading GreenBasket and a description field describing a London-based online marketplace connecting urban consumers with sustainable local producers.
One question at a time. No settings panel, no model picker, no configuration. The project brief is the interface.

Then, and only then

DesignX knowledge-base panel showing uploaded project files parsed and ready, with the artefact counter beginning to climb.
Knowledge base loaded and parsed. The artefact counter starts at 2 and ends the run at 134.
06

Control

Autonomy is a dial, not a switch.

Trusting a pipeline is not all-or-nothing. Four run modes move the same pipeline along a single axis: how often it stops and asks you. The gates never disappear; they only change who opens them.

Run mode

Gate behaviour

Phase boundaries pause

When you reach for it: The default.

Filled = pauses for you · Hollow = clears itself

Control is only half of trust. The other half is proof.

08

The eval

A rubric that tells you where it is weak.

Generation is easy to fake and hard to trust. So the pipeline scores its own output against a designer-QA rubric across eight dimensions, and reports the low scores as loudly as the high ones. The GreenBasket run came back at 0.0/5 with 16 named gaps, every one resolved to a production spec before delivery.

Dimension

Score / 5

Content & copy5
Token discipline5
Visual hierarchy4
Accessibility4
Responsiveness4
Component structure4
Performance4
Interaction detail3

The lowest score is the point

Interaction detail scored 3. That is the number I care about most.

A tool that grades its own work at five out of five is not an eval, it is marketing. The rubric exists to produce a worklist, and a 3 is where the next run starts. Each gap arrives as a named, specified fix: not “improve the interactions”, but the component, the state, and the intended behaviour.

16 gaps found · 16 resolved to production specs · 0 blockers at the concept gate

09

The record

The deliverable is not the screen. It is the system that produces screens.

A one-shot tool hands you a picture. DesignX hands you the argument for every pixel in it, and the system that will produce the next hundred consistently.

Stage 01

Evidence

Every source in the knowledge base, indexed and addressable by ID.
Stage 02

Insights

Findings that cite evidence IDs and carry an explicit confidence level.
Stage 03

Design system

Tokens, components and states, produced before any screen is drawn.
Stage 04

Production specs

Screens specified against the system, each traceable to the insight beneath it.
A DesignX design-system artefact listing resolved tokens, component definitions and their states.
The design system is generated as an artefact in its own right, with resolved token references, not implied by the screens.
A DesignX artefact listing the design principles derived for the project, each tied back to research findings.
Principles come out of the research, not out of a template, and the screens are then argued against them.

All of which is an argument. Here is what came out the other end.

10

The output

From brief to a working app.

GreenBasket was a throwaway test project, a fictional London grocery marketplace, run through the pipeline to see what came out the other end. Not a mockup: a working front end, built from the generated design system, with real copy and real product logic. These are its screens, untouched.

The GreenBasket homepage hero: a headline reading “As convenient as your usual shop. As meaningful as your local market,” body copy citing 127 verified producers across four London boroughs, a “See what's in season” button, and a photograph of a market produce stall.
Positioning, not filler. The pipeline wrote a claim it could substantiate: 127 verified producers across four named boroughs: because the knowledge base gave it the numbers to be specific with.

Three decisions worth pointing at

Every claim has a “Learn more”

Four trust claims sit above the fold, and not one of them is a bare assertion. Each opens its own methodology. “No dark patterns. Ever.” is an unusual thing for a generated store to promise, and it traces directly to the INS-08 implication.

A row of four trust cards on the GreenBasket homepage: 127 verified producers, Impact Score, Allergen data, and “No dark patterns. Ever.”: each with explanatory copy and a Learn more link.

Substitution is a preference, not a surprise

Every basket line carries its own “if unavailable” choice: substitute, contact me first, or remove. Fresh-grocery delivery lives or dies on this interaction, and the pipeline surfaced it per item rather than burying one global setting in an account page.

The GreenBasket checkout basket: two line items, each with an “If unavailable” row offering Substitute with similar, Contact me first, or Remove if unavailable.
The GreenBasket product grid: six product cards each showing an Impact Score badge, producer name, weight, price, and: where relevant: an allergen warning such as “Contains: Gluten (wheat), Sesame”.

Allergens on the card, not the detail page

Products are checked against the 14 major EU allergens, and the warning renders on the card itself. The Impact Score badge sits in the same corner on every tile: the kind of consistency that comes from specifying a component before drawing screens.

A block of four checkout reassurances: No hidden fees, 14-day cancellation right, Payment used only for this order, and No dark patterns. Ever.
At checkout, the same four commitments reappear, including the UK distance-selling cancellation right, stated in plain language.
The GreenBasket order confirmation screen, showing the order reference, delivery window, and itemised basket.
The run ends where a real one would: a confirmation screen, not a hero shot.
11

Proof

Proof it wasn’t a prototype.

We shipped DesignX to npm to test one thing: would anyone want design run as a governed pipeline? They did. Then we took it private once the idea was proven.

Public run

0

downloads in the first year

Releasesv1.0.0 → v5.5.3
52
Major versionsin ~3 weeks
4
Peak daylaunch
3,237

One proof run, end to end

Knowledge base in
16 KB
Artefacts out
134
Concept gate0 blockers
14 / 14
Designer-QA score16 gaps, all resolved
4.1 / 5
Total model spend
$5.78

The lasting idea isn’t the agents. It’s that when design becomes a pipeline, the deliverable stops being a screen.

It becomes the system that produces screens, one where every decision can be traced, defended, and re-run. That is the difference between a tool that makes design faster and one a senior team can actually stand behind in a review.