To measure if AI agents use your design system you let an agent build real screens with your component library and then grade the code it wrote. Open Design System Bench by Christoph Hellmuth from NordVPN does exactly that. A free static audit scores your repo in seconds. A behavioral benchmark grades real agent output from 0 to 100.
Christoph Hellmuth is Design Engineering Manager for Design Systems at NordVPN. His team builds Aurora. That is the design system behind NordVPN and other Nord Security products.
At our AI Jam on September 29 2026 he showed Open Design System Bench. These are my notes from his session plus the steps to run it on your own design system.

What you get from this post
- The numbers: what the benchmark found in Aurora and in eight open source design systems
- The fix: the one README sentence that cut hand-rolled components from 72% to 22%
- The setup: five commands to run the benchmark on your design system or one prompt for Claude Code
I shipped an AGENTS.md so why does the agent still build its own buttons?
Christoph's team did what everyone recommends. Aurora ships an AGENTS.md. Then they measured what the models actually did with it.
In 41 of 57 guided generations the models hand-rolled buttons, dialogs and toggles instead of using the design system they were told about. That is 72%.
"It just invented a button, and honestly, we were quite shocked, because we were thinking, like, yeah, best practices, they work well."

Every model launch comes with a scoreboard like SWE-bench or Terminal-Bench. None of them checks whether an agent uses your Button instead of inventing one.
That gap hurts when leadership asks what the 50 to 80 hours of AI readiness work brought. Without a number you only have a feeling.
"Cool, so can you prove it? And you… well, we did all these best practices, but not really, and that's not really what leadership wants to hear."
If you want the background on why agents break design systems read Your Design System Is Not Ready for AI Agents.
I want a to-do list for AI readiness in 10 seconds
Open Design System Bench has two tools. The first one is free and gives you the list.
| Static audit | Behavioral benchmark | |
|---|---|---|
| Question | What does your repo offer an agent? | What do agents actually write with it? |
| Checks | AGENTS.md, llms.txt, skills and catalog quality | Real screens built by coding agents in a fresh workspace |
| Time | 5 to 10 seconds in the live demo | Minutes to hours |
| Cost | Free with no LLM and no API key | Real LLM spend |
| Output | Score plus a list of what is missing | Score from 0 to 100 on six graded dimensions |
"In this just static check of 10 seconds, it also gives you basically a to-do list, and you can now go into Jira, create the tickets for the next month."

You can also run the audit against several systems at once. Christoph compared eight open source systems. These are the surface scores from his slide:
| Design system | Audit score |
|---|---|
| Chakra UI | 70.5 |
| Mantine | 70.1 |
| MUI | 66.3 |
| Ant Design | 66.0 |
| Primer | 62.2 |
| Carbon | 60.9 |
| Base UI | 59.5 |
| Radix | 34.1* |
*The extraction read 0 exports from the Radix umbrella package. Christoph's slide calls that a harness gap and not a grade on Radix.

The audit score is always the optimistic one. Ticking checkboxes is easy. The real score shows up once an agent has to use your system.
How do you write a benchmark task that does not give away the answer?
The benchmark ships with 10 tasks. One example is a password reset screen. Another is the area where a user cancels a subscription.
- Describe intent: a task never names a component. Otherwise you only test whether the model can read.
- Weigh the rubric yourself: in Christoph's example a red destructive cancel button counts for more than a support link.
- Let the linter catch leaks: a built-in linter checks your task before you run it. In his demo it failed a task because the prompt mentioned the real export
PasswordInput.

What exactly gets graded?
Each output is scored on six dimensions. The weights come from the README:
| Dimension | Weight | What it checks |
|---|---|---|
| imports | 10% | Only your package, React and local files are imported |
| apiFidelity | 25% | No hallucinated components or invented props |
| tokenDiscipline | 15% | No raw hex values, arbitrary Tailwind or hardcoded styles |
| a11yStatic | 10% | Accessible names, labels and valid ARIA |
| compile | 10% | TypeScript compiles |
| judgment | 30% | Per-task rubric graded by a separate model |
By default every task runs 3 times per model. The worst run counts so the score never looks better than what a designer would get on a bad day.

The hallucinations Christoph found in Aurora were very specific:
- Aurora has a
headingprop and the model invented aHeadingcomponent - Aurora has
Textand the model usedTypography - The model invented spacing props on
BoxandCard
"They hand over the AI generated prototype to developers. And they tell you, well, we don't really have any of these props, or this is an invented component."
Why did the repo read 20 points better than it behaved?
Christoph ran the benchmark on Aurora AppKit. The audit gave it 64.8. Six models building with it scored 44.9.

Only 10 of 55 guided cells compiled against the published types. The npm run produced about 250 type errors. About 100 of them came from direction="column" on a prop that only takes 'vertical' | 'horizontal'.
The uncomfortable part: their own AGENTS.md and DESIGN.md mentioned a cn helper and a bare Radix component that never existed in the codebase. Their docs were teaching the agent some of the traps.

Which sentence should I add to my README?
Their AGENTS.md told the model what not to use. Then they added one sentence to the README saying the package is installed.
| Setup | Used the design system | Hand-rolled |
|---|---|---|
| No guidance | 0 of 113 | |
| AGENTS.md (run against source) | 16 of 57 | 72% |
| AGENTS.md plus "the package is installed" (run against npm) | 43 of 55 | 22% |
The slide notes that the npm run also changed how the package was consumed. Christoph calls it a strong hint and not a controlled experiment.
"Even things you thought, oh, this is useless, like, why should I tell the AI that my design system package is installed? No, it actually makes a difference."

After fixing the findings the audit score went from 64.8 to 82.0. The fixes: export the types nobody had exported, generate llms.txt and a DTCG tokens.json from code, add a positive mandate plus a vocabulary crosswalk and delete the phantom docs.
| Audit dimension | Before | After |
|---|---|---|
| Overall surface | 64.8 | 82.0 |
| Export hygiene | 65 | 100 |
| Docs greppability | 69.3 | 100 |
| Enablement surface | 55 | 90 |
| Token machine-readability | 66.6 | 81.6 |
| Catalog quality | 86.4 | 88.1 |
| Vocabulary | 61.4 | 61.4 |
| Deprecation | 40 | 40 |
Vocabulary and Deprecation stayed put on purpose. Renaming the API to match a model's guess is a breaking change so the team documented the mapping instead. Nothing in Aurora is deprecated yet.

He also tested models with and without instruction files. In the session he said adding one file moved the scores by 20 to 30 points. His write-up lists these mean scores:
| Model | No guidance | Instruction files, skills and catalog |
|---|---|---|
| DeepSeek V4 Pro | 58.7 | 84.8 |
| Sonnet 5 (single-shot) | 51.2 | 75.3 |
How do I run Open Design System Bench on my design system?
The repo supports React and TypeScript libraries plus web components built with Stencil or Lit. Your system can be local source or an npm package.
- Clone open-design-system-bench
- Install and set it up:
npm install
npx tsx src/cli.ts init # wizard asks for package name, location and how you serve it
npx tsx src/cli.ts doctor # checks your config and CLI setup
npx tsx src/cli.ts extract # builds the component catalog
npx tsx src/cli.ts run --profile smoke # 2 cells in about 5 minutes
- Run the bigger profiles once the smoke test passes.
smallhas 5 cells for weekly checks andfullhas 90 cells for a quarterly baseline.

You do not have to type any of this. Christoph and I both said it in the session: paste the repo link into Claude Code and ask it to install the benchmark and run it against your design system.
Install https://github.com/christophhdesign/open-design-system-bench and run it against my design system
| Run | Time | Cost |
|---|---|---|
| Static audit | Seconds | Free |
| Smoke profile | About 5 minutes | |
| Full run with Claude Code | Half an hour to an hour | |
| Run with 6 models | About 3 hours | |
| Run with the top models through the API | A few hundred dollars |
A subscription is much cheaper than the API.
Christoph's tip for the report: give it to Claude and ask for a PDF for leadership with an executive summary, the models tested, the results and the costs.
Can I score my Figma file too?
Vee asked in the Q&A whether you can score Figma and code together. Christoph's answer: a Figma file alone mostly tests how well the AI uses the Figma MCP. It says little about your design system.
His recommendation for the Figma to code workflow is Code Connect.
"Honestly, at least from what we've seen, CodeConnect is totally worth having for, your Figma to code, workflow."
To test the Figma side of your system on its own see Tutorial: How to test a Figma design system with Claude Code.
What I'm taking back to my own work
These are my own takeaways and not Christoph's:
- Run the free audit first. It takes seconds and gives you a ticket list.
- Never trust the audit score alone. Aurora lost 20 points once agents had to build with it.
- Tell the agent what to use and not only what to avoid.
- Grep your own AGENTS.md for helpers and components that do not exist.
- Write benchmark tasks the way a designer would brief a screen. Never name the component.
How other teams make their systems readable for agents: Miro with MCP and Claude Code skills and Spotify with Encore. Christoph and his hackathon team built a Figma plugin in 48 hours with AI: How Designers built a working Figma Plugin.
Links from the session
- Open Design System Bench on GitHub
- Christoph's write-up with all results
- Open Design System MCP: an MCP server from Christoph that exposes your real components and tokens to agents so they stop inventing names and props
- What Does AI-Readiness Mean in Design Systems? by Cristian Morales
- Design Engineer job in the Aurora team at Nord Security
Watch Christoph's full session and the live demo
The recording of this AI Jam is part of the AI Design Systems Conference 2027 ticket. The conference runs online on March 10 and 11 2027 with speakers from Anthropic, OpenAI, Atlassian, Netflix and Yahoo.
📐 AI Design Systems Conference 2027
- ✅ Live access on both days with Q&A
- ✅ All conference recordings with lifetime access
- ✅ Access to AI Jam recordings including this session
- ✅ Templates, files and prompts from the speakers
👉 Get your ticket at the $499 launch price. After October 20 the price goes up


