Design system, AI platform — Nayara Marques
Design systemAI platform · private capital markets2026

A design system AI can build from

Most new screens on the platform were built with AI tools, and the design system gave those tools two different answers: one in the design tool, another in the code. I moved the single source into Claude Design, where the screens are built. I wrote eight families of rules so AI tools reuse existing components instead of guessing, and set up a way for anything new an AI tool draws to reach me for review.

AI tools built most new screens, and the design system gave them two answers. I moved the single source into Claude Design and wrote rules so AI tools reuse components instead of guessing.

Role
Sole designer. Tokens, components, rules, templates, governance.
Users
Engineers building UI, AI tools generating it, one designer reviewing it.
Scope
Three products and a marketing site, one codebase. One designer, six people building against the system.
Built in
The design system lives in Claude Design, the build environment, as the single source. Engineering receives it through Claude Code, as a set of rules their tools can read.

01 — Context

The products: an AI toolkit for investment workflows, a network for industry groups, and a marketplace where allocators source and diligence investments.

Because I was the only designer, the system had to work as a set of instructions, precise enough that whoever builds next, an engineer or an AI tool, gets the answer I would give.

Figures are wireframes of the system's structure; names and values are as specified.

Scope

  • 48components now, each with a contract and a spec card
  • 8rule families, plus a decision map and one traced screen
  • 9templates as sanctioned starting points

02 — The problem

The design system had two problems. Its two copies, one in the design tool and one in the code, no longer matched. And even where components existed, screens still came out inconsistent.

The component library in the design tool was complete. The one in the codebase, which is what actually shipped, was incomplete and inconsistent. So every question about a colour, a variant or a hover state had two answers, and engineers had to ask me which was right. Forty-odd components existed, and screens still came out inconsistent, because nothing said how to use them together. No rule said whether a list page has a white or a grey background, how titles are sized, which of four tab-shaped components is for navigation, or when a shadow is allowed.

What the audit found

Two font weights with no font file, so the browser faked them. Two shadow scales with near-identical names and different values. One type name meaning 14px in one scale and 18.7px in the other. title declarations doing one job, in a single product tab families, four different bottom edges With AI tools, these problems appear faster. An AI tool draws a plausible table or banner in seconds. Because it looks finished, the team treats it as a decided design, and a reviewer cannot tell it was invented without opening the code.

03 — Constraints

The system had to meet four conditions: where its single source lived, how components enter it, how gaps are handled, and how one reviewer keeps up.

A design tool that holds only pictures of components has to be re-coded by hand before anyone can build with it. So the single source had to live in the tool where screens are built. I did not import a ready-made component kit. I rebuilt each component on our own tokens, reviewed it on its own and gave it a spec card: a short page with its job, where not to use it and what to use instead. AI tools fill anything unspecified with a plausible guess. So every rule ends in an exact value, and where no rule exists yet, the system says so in a published list. Review processes built for a design team wait on people a one-designer team doesn't have. Nothing could block a build or need two people to agree.

04 — Approach & key decisions

Four stages, in this order, from an audit of what shipped to a route for reporting what the system doesn't cover yet.

Every header, token and near-duplicate component catalogued against what shipped.I named Claude Design as the canonical source. The code repository is still read for how components are built, but it no longer decides.Rules for using components together, such as which background a list page has, each written to pass or fail. Section 06 lists them.A way for engineers and AI tools to report any part they draw that the component catalogue does not cover.
I moved the design system into Claude Design, where screens are built, and the team stopped working from a design tool that holds only pictures. The codebase is still read for how each component is built, but on any conflict of colour, type, naming or behaviour, the canonical source wins. The rule that the canonical source wins is the first line of the design system's readme, the page every builder opens first, and of its rule set. A tiebreak only works if nobody can misread it. A layer above the components holds the decisions a component's settings cannot express, such as when a shadow is allowed (section 06). It also has a decision map from each job to its component, and one reference screen where every decision is traced to its rule. When a screen needs a part the system lacks, people and AI tools draw it anyway, and forbidding that only hides it. So whoever draws a new part tags it in the prototype and adds a row to the register, a shared list, with the part's job and why each existing component didn't fit. Building continues. The reviewer then makes the part a component, rejects it and names the component to use instead, or holds it until it appears again.

05 — What I built

The system has four layers, each built on the one below it, and a register for anything none of the layers covers yet.

The four layers are tokens, components, applied rules (the rules layer from section 04) and templates. Each answers a question the one below leaves open. Most of the inconsistency came from how components were combined, which no rule covered; most design systems leave this layer to design review.

Layer 01Tokensfix the values

Leaves open: which token, where.

Layer 02Componentsfix the parts, one contract each

Leaves open: which one for this job, and how they sit together.

Layer 03Applied rulesfix the decisions between parts

Surfaces · titles · geometry · spacing · elevation · states · motion · limits. Pass or fail, ending in a self-check.

Layer 04Templatesfix where a page starts

Leaves open: only the content.

Register

For a part no layer covers. Building continues while it waits for review.

Drawn and tagged
↓
Row filed: its job, and why each existing part was wrong
↓
Reviewed
Becomes a component → joins Layer 02
Rejected, with the replacement named
Held until it recurs
The four layers and the register. Whoever builds a screen works up from the bottom layer to the top. A part no layer covers goes to the register, and the reviewer's verdict can turn it into a new component.

Decision map

Most inconsistent screens did not invent anything new. They used a reasonable component for the wrong job: a filter drawn as navigation, a status pill used as a count, a hand-built menu. The decision map lists each job, the right component for it, and the plausible substitute people kept choosing.

Decision map · job → component → wrong choice
The job Use Reached for instead
Move between a record's sub-pages Underline tabs The filled segmented switcher, which means a view and not a route
Narrow a list to this quarter Toggle group in the toolbar A second tab row: a filter drawn as navigation
Open a status or download menu Menu button Button plus chevron plus a hand-built surface
Show a count beside a nav row The rail's built-in numeral A badge, which in this system means a status
Lay cards out in columns Card grid A hand-rolled grid, which gives every card its own height
Five rows from the decision map. Each row names a job, the component to use and the substitute to avoid. Naming the wrong answer changed what engineers and AI tools built more than naming the right one.

Register

When an engineer or an AI tool needs a part the catalogue lacks, the register is how it reaches me. If the same part turns up in two projects, it stays one row with two sightings, so I can see which new parts keep coming back.

Nothing here is canon until it ships
01 Drawn

Tagged in the prototype. The build keeps going.

02 Reported

Picture, spec, and what was rejected and why.

03 Reviewed
Approved to buildRejected, with the name to use

One reviewer. A rejection has to name the replacement.

04 In the system

Rebuilt on tokens with a spec card. The hand-drawn markup is deleted.

Component

Seen twice · not composable · tokens only.

Variant

A different meaning, not a different look.

Composition

Same arrangement twice · carries decisions · shorter than the markup it replaces.

Four states and three tests. A new part moves from drawn, to reported, to reviewed, to in the system, while building continues. If the reviewer approves a part, three tests decide what it becomes: a new component (seen twice, can't be built from existing parts, uses only tokens), a variant of an existing one (it means something different, not just looks different), or a composition of existing parts (the same arrangement seen twice).

AI use while exploring

Most ideas on this team start with an AI tool exploring options, which is where invented parts first appear.

That is why each component's spec card states its job, where not to use it and what to use instead. An AI agent reads this, reuses the component that fits, and tags anything it has to invent, so the new part reaches the register for review.

An agent matches a need to an existing part instead of redrawing it. New parts are reported, then become components or are replaced by existing ones.

The AI features of the platform joined the system this way. The composer (the bar where users type to the AI), edit mode (where the AI proposes changes to a document) and the review step (where the user accepts or rejects them) were reported as new parts, then became components with rules. See AI surfaces.

06 — The rules

Eight rule families, covering the decisions between components. Every answer is a number, a token or a named component.

The test: could an engineer who has never spoken to me apply it, and a reviewer check it without the design file? “Use generous spacing” fails. “Twenty-four for sections, cards and every inset” passes. People and AI tools need rules this exact for the same reason: neither can ask me what I would prefer every time.

I took every value in the rules from the components themselves, never from memory, because a rules document that disagrees with the components would be a third source of truth.

Family The rule
SurfacesPage, then card, then content. Never a third fill. Floating layers are lighter than what sits under them, and one grey is hover only.
TitlesSize is set by the container, never by importance. One size per container type, listed.
Body and weightOne default for body, cells and field values. Four weights with a named job each, and two deleted because no font file backed them.
GeometryFour legal control heights and six legal radii. Nothing else, including “close enough”.
SpacingA short ladder, and one number for sections, cards and every inset, so page rhythm cannot be renegotiated per screen.
ElevationNo shadow by default. A shadow means the thing floats. It never signals importance.
StatesEight required states per component, each with one treatment. Every list owes an empty state and a loading state.
MotionDurations by interaction type, two easings, and nothing animates layout.

Self-check and override

The rules end in a short checklist that whoever builds runs through: colours only from tokens, never typed in by hand; corner radii, control heights, font sizes and spacing only from the allowed values; tables, buttons, inputs and selects only as system components; shadows only on floating layers. If applying a rule needs interpretation, the builder brings it to me as a question instead of guessing.

When engineers adjust a component on a page, they may change its layout. They may not change anything that defines what the element is: background, colour, border, radius, height, padding or type. If they need a version the component doesn't have, they ask for a new variant instead of restyling it. This rule stopped engineers restyling a component inside a single page, which had been creating unofficial copies of it.

For an AI model, exact rules also save work: it invents fewer parts, and fewer screens have to be generated a second time.

Published gaps

The rules come with two published lists. The undecided list (page width, grid, breakpoints, layering, the success, warning and info colours, validation timing, table density) tells a builder who needs one of these to ask first instead of inventing an answer. The not built list, from a mobile nav drawer to charts, tells them to report the gap instead of shipping a lookalike.

Publishing these gaps was the least intuitive decision, and it changed the most. When the system does not cover something, a close approximation looks like an official decision; a published gap tells people to ask instead.

07 — Outcome

The system is in use and still growing. Three things have changed for the engineers and AI tools that build with it.

Engineers and AI tools build against the rules without a designer reviewing each screen. Decisions that used to need my judgement are now written as exact values. Anything drawn outside the catalogue arrives as a row in the register. I review a list of new parts instead of reading every screen's code, which is the only way one reviewer can keep up. The undecided and not-built lists ship with the rules. Someone who needs breakpoints or a chart asks me first, instead of inventing an answer that has to be undone later.

Nothing is measured yet. The counts at the top are scope, not evidence.

08 — What this case doesn't cover

These are the parts of the work this case leaves out, and why.

Not measured yet. The cost of generating each screen, counted in the AI model's usage tokens (not design tokens), should fall as gaps close; it is not tracked yet. One canonical source left the repo knowingly behind. The re-sync has not happened. Each of the forty-eight components has its own decisions, states and spec card. This case covers the layers and rules that connect them. Dark-theme tokens exist, but no rule says how to use them. Accessibility is not covered in this case. Charts, a drawer and an accordion are named gaps. Six active users, one reviewer. The system assumes local decisions win and someone files the rows. Both are deliberate, and untested with more designers and projects.
Design systems

Let’s make a system your team and its AI tools can build from.

Tokens and components are the easy half. I write the rules for combining them, so AI tools reuse what exists, and I set up the review that catches what they invent.

Let’s talk