A design system AI can build from
Most new screens on the platform were built with AI tools, and the design system gave those tools two different answers: one in the design tool, another in the code. I moved the single source into Claude Design, where the screens are built. I wrote eight families of rules so AI tools reuse existing components instead of guessing, and set up a way for anything new an AI tool draws to reach me for review.
AI tools built most new screens, and the design system gave them two answers. I moved the single source into Claude Design and wrote rules so AI tools reuse components instead of guessing.
- Role
- Sole designer. Tokens, components, rules, templates, governance.
- Users
- Engineers building UI, AI tools generating it, one designer reviewing it.
- Scope
- Three products and a marketing site, one codebase. One designer, six people building against the system.
- Built in
- The design system lives in Claude Design, the build environment, as the single source. Engineering receives it through Claude Code, as a set of rules their tools can read.
01 — Context
The products: an AI toolkit for investment workflows, a network for industry groups, and a marketplace where allocators source and diligence investments.
Because I was the only designer, the system had to work as a set of instructions, precise enough that whoever builds next, an engineer or an AI tool, gets the answer I would give.
Figures are wireframes of the system's structure; names and values are as specified.
Scope
- 48components now, each with a contract and a spec card
- 8rule families, plus a decision map and one traced screen
- 9templates as sanctioned starting points
02 — The problem
The design system had two problems. Its two copies, one in the design tool and one in the code, no longer matched. And even where components existed, screens still came out inconsistent.
What the audit found
03 — Constraints
The system had to meet four conditions: where its single source lived, how components enter it, how gaps are handled, and how one reviewer keeps up.
04 — Approach & key decisions
Four stages, in this order, from an audit of what shipped to a route for reporting what the system doesn't cover yet.
05 — What I built
The system has four layers, each built on the one below it, and a register for anything none of the layers covers yet.
The four layers are tokens, components, applied rules (the rules layer from section 04) and templates. Each answers a question the one below leaves open. Most of the inconsistency came from how components were combined, which no rule covered; most design systems leave this layer to design review.
Leaves open: which token, where.
Leaves open: which one for this job, and how they sit together.
Surfaces · titles · geometry · spacing · elevation · states · motion · limits. Pass or fail, ending in a self-check.
Leaves open: only the content.
For a part no layer covers. Building continues while it waits for review.
Decision map
Most inconsistent screens did not invent anything new. They used a reasonable component for the wrong job: a filter drawn as navigation, a status pill used as a count, a hand-built menu. The decision map lists each job, the right component for it, and the plausible substitute people kept choosing.
| The job | Use | Reached for instead |
|---|---|---|
| Move between a record's sub-pages | Underline tabs | The filled segmented switcher, which means a view and not a route |
| Narrow a list to this quarter | Toggle group in the toolbar | A second tab row: a filter drawn as navigation |
| Open a status or download menu | Menu button | Button plus chevron plus a hand-built surface |
| Show a count beside a nav row | The rail's built-in numeral | A badge, which in this system means a status |
| Lay cards out in columns | Card grid | A hand-rolled grid, which gives every card its own height |
Register
When an engineer or an AI tool needs a part the catalogue lacks, the register is how it reaches me. If the same part turns up in two projects, it stays one row with two sightings, so I can see which new parts keep coming back.
Tagged in the prototype. The build keeps going.
Picture, spec, and what was rejected and why.
One reviewer. A rejection has to name the replacement.
Rebuilt on tokens with a spec card. The hand-drawn markup is deleted.
Seen twice · not composable · tokens only.
A different meaning, not a different look.
Same arrangement twice · carries decisions · shorter than the markup it replaces.
AI use while exploring
Most ideas on this team start with an AI tool exploring options, which is where invented parts first appear.
That is why each component's spec card states its job, where not to use it and what to use instead. An AI agent reads this, reuses the component that fits, and tags anything it has to invent, so the new part reaches the register for review.
The AI features of the platform joined the system this way. The composer (the bar where users type to the AI), edit mode (where the AI proposes changes to a document) and the review step (where the user accepts or rejects them) were reported as new parts, then became components with rules. See AI surfaces.
06 — The rules
Eight rule families, covering the decisions between components. Every answer is a number, a token or a named component.
The test: could an engineer who has never spoken to me apply it, and a reviewer check it without the design file? “Use generous spacing” fails. “Twenty-four for sections, cards and every inset” passes. People and AI tools need rules this exact for the same reason: neither can ask me what I would prefer every time.
I took every value in the rules from the components themselves, never from memory, because a rules document that disagrees with the components would be a third source of truth.
| Family | The rule |
|---|---|
| Surfaces | Page, then card, then content. Never a third fill. Floating layers are lighter than what sits under them, and one grey is hover only. |
| Titles | Size is set by the container, never by importance. One size per container type, listed. |
| Body and weight | One default for body, cells and field values. Four weights with a named job each, and two deleted because no font file backed them. |
| Geometry | Four legal control heights and six legal radii. Nothing else, including “close enough”. |
| Spacing | A short ladder, and one number for sections, cards and every inset, so page rhythm cannot be renegotiated per screen. |
| Elevation | No shadow by default. A shadow means the thing floats. It never signals importance. |
| States | Eight required states per component, each with one treatment. Every list owes an empty state and a loading state. |
| Motion | Durations by interaction type, two easings, and nothing animates layout. |
Self-check and override
The rules end in a short checklist that whoever builds runs through: colours only from tokens, never typed in by hand; corner radii, control heights, font sizes and spacing only from the allowed values; tables, buttons, inputs and selects only as system components; shadows only on floating layers. If applying a rule needs interpretation, the builder brings it to me as a question instead of guessing.
When engineers adjust a component on a page, they may change its layout. They may not change anything that defines what the element is: background, colour, border, radius, height, padding or type. If they need a version the component doesn't have, they ask for a new variant instead of restyling it. This rule stopped engineers restyling a component inside a single page, which had been creating unofficial copies of it.
For an AI model, exact rules also save work: it invents fewer parts, and fewer screens have to be generated a second time.
Published gaps
The rules come with two published lists. The undecided list (page width, grid, breakpoints, layering, the success, warning and info colours, validation timing, table density) tells a builder who needs one of these to ask first instead of inventing an answer. The not built list, from a mobile nav drawer to charts, tells them to report the gap instead of shipping a lookalike.
Publishing these gaps was the least intuitive decision, and it changed the most. When the system does not cover something, a close approximation looks like an official decision; a published gap tells people to ask instead.
07 — Outcome
The system is in use and still growing. Three things have changed for the engineers and AI tools that build with it.
Nothing is measured yet. The counts at the top are scope, not evidence.
08 — What this case doesn't cover
These are the parts of the work this case leaves out, and why.
Let’s make a system your team and its AI tools can build from.
Tokens and components are the easy half. I write the rules for combining them, so AI tools reuse what exists, and I set up the review that catches what they invent.