Independent test · single operator · August 2026

What survives when a
design system meets Omma

One prompt, held constant. One variable changed at a time. Every generated project measured with a script rather than read by eye. This is what happened when a real, published design system was fed into the canvas — where it transferred completely, where it stopped, and what one sentence in a prompt turned out to be worth.

Naming fidelity, with spec
0%

Sixty tokens declared, sixty carrying the system's own names. Zero generic substitutes, including two private namespaces a generator would normally normalise away.

Compact file vs full spec
×0

A 5.6 KB instruction file matched a 103 KB specification on naming and beat it on value fidelity. Eighteen times less input, same transfer.

Lines needed to bind the 3D layer
0

One sentence of intent moved the scene from hand-copied values to live resolution of the page's custom properties. A thirty-line contract was not required.

Where this started

Not with a protocol. With an ordinary afternoon of trying to reproduce what had been demoed — three weeks before any of the measurements below, no design system loaded, no method, no instrument.

The whole prompt:

a landing page hero containing a 3d brain made of tiny opalescent cubes

Two iterations. The first returned a blank canvas and Cannot read properties of null (reading ‘appendChild’). The second rendered: a soft white cloud, with no anatomy, no cubes visible at any scale, and no opalescence.

What the agent reported having built, in the message attached to that same render: roughly nine thousand instanced cubes arranged into an anatomically-inspired lobed surface with a central fissure, gently rotating with mouse parallax, an iridescent material using physical transmission, and bloom post-processing for the glossy finish.

Read the description and the render side by side. The gap between what was reported and what appeared is the first datapoint of this entire exercise — and it arrived weeks before anything existed to measure it with. Every structured finding further down is, in one form or another, that same gap held still long enough to be counted.

Second attempt: the specification loaded, and thirteen rounds

Same subject, this time with the design system loaded and a fuller prompt:

a landing page hero section containing a 3d brain slowly rotating made of tiny
opalescent cubes like a repeater, while hovering the cubes scale

This one works. It renders a plausible product page — navigation, an eyebrow chip, a two-line headline with the accent on the second line, a lede, a call to action, and a cube-built brain below the fold. The model invented a brand and wrote the copy for it, unprompted and coherently.

Then thirteen rounds of iteration. Every change requested — including changes that should have been structural — produced a minimal visible delta. The rounds were not idle: they contained real work, including a correct diagnosis of a crash in the hover maths (a degenerate ray-plane intersection returning null every frame, replaced with a direct camera-ray unprojection). The engine responded. The result barely moved.

Third attempt: a clean session, two rounds

Same specification, same subject, a shorter prompt, and one deliberate change: a brand-new chat, with none of the accumulated context.

design a 3d opalescent 3 brain made with cubes, interactive on hover

Two rounds. The result is the best of the three: cubes actually legible as cubes, at varying sizes, with a genuine iridescent material shifting between the system's magenta and cyan, and a hover interaction that reads — nearby cubes ripple outward and the glow lifts. The framing around it came out in the system's own register, down to a mono uppercase figure label.

It is still not a brain. Three attempts in, the anatomical form never arrived, and the request to make the cubes vary in size and resemble a brain more closely had to be made explicitly here too. But the direction of travel is unmistakable: two rounds in a clean session beat thirteen rounds in a loaded one.

One note, unmeasured: in the third attempt the model stated the result stayed inside the loaded token system. The visible palette is consistent with that — an iridescence travelling between the system's two accent colours — but no code was exported and nothing here was checked.

What the three attempts say together

The hard part is not the first prompt. It is the second. Getting something on screen is easy and often good. Moving it towards what you had in mind is where the effort goes — and thirteen rounds can end close to where they started, while two rounds from a clean slate go further. For a user without the patience or the technical footing to open the engine, that is where the trial ends: not at a failure, at a plateau.

Read strictly, the fastest repair for a session that has drifted was to abandon it. That is worth sitting with, because it inverts what a conversation is supposed to be for. If accumulated context makes steering harder rather than easier, the session's history is a liability, and the leverage is not in a better first generation nor in a better editing loop — it is in collecting the intent before the first generation happens. A few structured questions instead of an open field.

Everything measured in the rest of this document is a version of the same idea, arrived at from the other end: what the runtime knows before it starts determines how much of the work is left to iteration. One sentence supplied up front bought a runtime binding that no amount of iteration had produced.

This arm is not measured

Neither attempt was exported and the checker did not exist yet. Both are reported as observations, not measurements, and no number in this document depends on either.

They earn their place for two reasons. This is where a trial user's experience actually begins — not at the successful run, at the one that goes nowhere and at the one that goes nowhere slowly. And roughly three quarters of all the credits spent across this work went into this unstructured phase, against a quarter for every measured arm that followed. The cost of not having a method was about three times the cost of having one.

Method

Nothing below is an impression.

One canonical prompt, frozen at the first run and never edited. It describes a screen, not an aesthetic: no colours, no fonts, no measurements. Changing it would break the comparison, so it did not change.

One variable per run. Between two arms, only the specification loaded changes, or only one clause added to the prompt. Each arm runs in a fresh project.

Code is the evidence. A screenshot tells you whether it is beautiful; only the code tells you whether the token arrived. Every run was exported and analysed.

Six dimensions, measured: token purity, naming fidelity, value fidelity, component architecture, motion discipline, and how far the token system reaches into the 3D layer — plus one metric added for this test, described below.

The scale was anchored before the runs, not after. A floor: a generic hero built with no specification at all. A ceiling: the design system's own source. The ceiling matters — 100% does not exist even there, and an output read against perfection is read wrong.

What transferred

With the specification loaded, the system reaches the DOM and the stylesheet essentially intact. The two private, non-obvious namespaces survived with their own names — the case where a generator normally reaches for --gray-* and --primary.

Dimension Floor · no spec Baseline · no spec With spec Ceiling · the system's own source
Naming fidelity 0% ~25% 100% 96.8%
Token purity 7.4% 22.8% 70.3% 74.3%
Value fidelity 0% 0% 96.3% 68.3%
Canonical component classes 1 / 15 2 / 15 7 / 15 13 / 15
Off-system easing curves 1 / 1 — 0 / 4 0 / 2
Token reach into the 3D layer 0% 0% 0% —

The last row is the interesting one, and the rest of this document is about it.

A compact file matched a large one

The same prompt was run against the full 103 KB specification and against a 5.6 KB instruction file derived from the same system. Naming fidelity: 100% both times. Value fidelity: 96.3% for the large file, 100% for the compact one. Token purity within two points. The compact file carried fewer component classes — five against seven — which is the one place the extra size paid for itself.

Read as a product fact: the runtime does not need a large document to honour a design system. It needs the right small one.

With one qualification, added after the fact and worth stating plainly: that comparison measures the generated code, and it holds there because the compact file inlines its token values as text and the model reads text. One level up, in the token model the platform builds for itself from the same file, the two inputs are not equivalent at all — the parser observation below shows what it makes of the prose, and shows that the generation figures in this table are unaffected by it. A compact file is the right shape; compact prose is the right shape for one consumer and the wrong shape for the other.

Repeated one week later

The identical prompt with the identical file was run again seven days after the first measurement, in a fresh project. Naming fidelity 100% both times; both private namespaces surviving with the same counts, twelve and eight; token purity within six tenths of a point; the 3D layer behaving identically. On this arm the generation behaviour is unchanged over the week.

The metric this test had to add

No existing measurement covered it, because no previous target had a layer the stylesheet cannot reach. Three states, and the distance between the second and the third is the whole subject:

State 01

None

The generator ignores the system outside the CSS. The 3D layer is invented from scratch.

State 02

Mirror

It recognises that the system should govern that layer, and hand-copies the values into JavaScript. Correct today, drifting tomorrow.

State 03

Live

It resolves the values at runtime from the page's own custom properties. Change a token in the stylesheet and the scene changes with it.

Without any 3D instruction, the default is mirror

With the specification loaded and nothing said about the 3D layer, the generator understood that the system was meant to govern the scene — it quoted the system's own colour-budget law to decide which single object earned an emissive surface — and then wrote a block of colour constants at the top of the scene file, annotated “token mirror, kept in sync”. Thirty-four values, all exact, zero drift.

The comment is the admission: it knows it is a copy, and it knows the copy will rot. The specification forbids precisely this. The rule was broken because no channel existed — not one line of the scene read a custom property. That is a wiring gap, not disobedience.

In a later run the same behaviour appeared without the annotation: the palette values simply inlined into the three.js calls, correct and untraceable. Same fidelity, no admission.

What each instruction bought

Five arms, same specification, same canonical prompt, each with a different amount of instruction about the 3D layer. Everything to the right of “binding” is the price paid for it.

What the prompt added Binding Naming fidelity Values copied by hand Reduced motion in the scene
Nothing mirror 100% 6 absent
A plain-language request naming the tokens mirror, reasoned — 15 absent, but offered
One sentence of intent: never copy a token value into JavaScript live, and reactive 90.4% 11 absent
The mechanism named: resolve via getComputedStyle live 74.2% 1 absent
Three clauses: no copies, read the system's own names, honour reduced motion live 100% 0 present

All four binding arms consumed the same input — the compact instruction file, unchanged across every run. The only variable between them is the clause printed in the appendix. The naming figures in that column are therefore comparable to each other, and the first row is the same file scoring 100% with nothing added.

One sentence was enough

Asking only for the outcome — never copy a token value, if a token changes the scene changes with it — produced a live resolver, a mutation observer that rebuilds the scene when the stylesheet or the colour scheme changes, a correct cubic-bézier solver that parses the curve out of the CSS string rather than retyping it as numbers, and a radius token driving the roundness of the geometry. None of that was asked for by name. The mechanism did not need teaching. Only the intent did.

The price is the vocabulary

Binding forces a choice the plain generation never had to make: which names do I read? And the generator chose its own. One run invented a seven-name semantic layer — --color-accent, --color-card, --color-border — while the design system already ships exactly those roles under its own names, none of which appear in the output. Another went further and added a twelve-name scene layer on top: fog, grid, particle, ring, and three scene durations.

Both layers are defined in terms of the system's real primitives, so the lineage is intact. But the scene is now bound to names the design system does not ship. Dropped into the real page, it would resolve nothing and fall silently back to the literals kept as defaults. Binding and naming are two primitives, not one — and today neither is the default.

The third arm closes it: with the naming clause added, naming fidelity returns to 100%, the binding stays live, and every hand-copied value disappears — including the fallbacks. That run also throws a typed error when a token is missing instead of substituting a literal, which was the third clause, and honours reduced motion inside the WebGL layer, which was the second.

Four reproducible observations

3 of 3 Agent file deletions do not persist

Three times, in three separate projects, a stray entry file at the project root broke the preview by loading an HTML file as a JavaScript module. Three times the agent diagnosed the cause correctly, in detail, and stated it had deleted the file. Three times the file was still there afterwards — once with the deletion claim on screen next to an export that still contained it, still carrying the exact line the message said had been removed.

Moving the entry files to the root worked immediately, both times it was tried. This reads as a scaffold behaviour, not a model failure: the reasoning is right and the delete is issued, and something downstream does not honour it. The cost lands on the user as rounds, credits, and the impression that the agent is not telling the truth.

8 and 9 rounds The error message does not name the file

Unexpected token ‘<’ arrives without a path. The cause was an entry file loading an HTML document as a JavaScript module — the 3D scene was correct and running the whole time. The agents rewrote it anyway, repeatedly, because nothing told them where to look.

Eight rounds in one project, nine in another, to get past an error that was never in the generated code. Pointed at the right file by hand, it resolved in a single shot for sixteen credits. And that figure is a floor rather than a ceiling: the later rounds were not the runtime working unaided.

Naming the file in the error is probably the highest value-to-effort change in the whole debug loop, because every wrong attempt is spent from the user's credit balance. The user pays for the ambiguity.

4 occurrences Reported work and the diff disagree

Beyond the deletions above, one round returned a six-section report describing a 3D refactor in specific technical detail — emissive intensity, light colours, easing sources — against a scene file that was byte-identical to the previous version.

The significant part is that the reasoning in that report was correct, and better than correct: the observation that shadow tokens encode blur and spread rather than a scalar distance, so there is no CSS property from which to read a depth, is true and not obvious. The model understood the contract. Then it reported having applied it. Asking for verbatim evidence — paste back the actual lines, not a summary — was what converted the next round from a report into a real refactor.

5 of 5 Reduced motion never reaches the WebGL layer

Across every run where it was not explicitly requested, there was no prefers-reduced-motion handling in the scene — with objects in perpetual rotation, orbiting rings and satellites. The CSS-level kill switch is emitted correctly and cannot reach WebGL. For a viewer with vestibular sensitivity, that hero cannot be turned off.

Nobody in this category appears to be handling it. Whoever does it first has a differentiator that enterprise buyers understand without explanation.

silent The design-system parser accepts prose and invents tokens from it

The same system was loaded twice, in two forms: once as a structured document with a machine-readable header, once as the compact prose instruction file. The style panel shows what the platform made of each.

From the structured document, the token model is correct: the brand and neutral colours by name, the full neutral ramp, and a spacing scale reading s-1 4px, s-2 8px, s-3 12px, s-4 16px and up.

From the prose, the extracted spacing scale reads:

utline    2px
offset    3px
labels   10px
adding   14px

These are not tokens. They are fragments of words. utline is the tail of outline; adding is the tail of padding; offset comes from outline-offset. Each carries a number scraped from the sentence it appeared in. The colour panel shows an unnamed row of dark swatches.

The mechanism is visible in the output itself: this is substring matching applied to prose. outline-offset: 2px yields a token called utline worth 2px. padding: 14px yields adding. A pattern written for a structured file, run over sentences, returns the tail of every word it half-matches.

The generation is unaffected, and it is worth showing why rather than asserting it. The compact file states in its own first lines that the token values are inlined so that it works with no other files — the text contains --obsidian-950 #07090B and the rest in full. The language model reads that text. The parser is a second, separate consumer of the same file, and only the second one fails.

The generated code carries the proof. In the compact arm: sixty-six custom properties declared, sixty-six of them carrying the system's own names, zero invented, zero off-scale durations, two raw pixel values in the whole file. Had the generator been working from the parsed model, --utline and --adding would appear in the output and a 14px step would appear in the spacing. Neither does.

No error is raised and no warning is shown. The panel presents a broken system exactly as it presents a correct one, and a user pasting text instructions has no way to know which of the two they are looking at. That is what makes this quiet: the output looks right while the platform's own model of the system is nonsense, so nothing in the experience ever surfaces the gap.

across sessions One three.js clock bug recurs

The animation loop calls the elapsed-time accessor before the delta accessor, which zeroes the delta — measured at 0.000024 radians instead of 0.302 after two and a half seconds. Anything driven by delta barely moves. It appeared in the first session, was corrected, reappeared in the next project, and appeared again a week later in a fresh project. It is not that it does not learn between projects: it does not learn between sessions.

What Omma does that its category does not

It tries to obey the design system inside the 3D layer. No other builder in this cluster attempts that at all. Given the specification and no 3D instruction, it reached for the system's own colour-budget law to decide which surface earned an emissive, and then went looking for a way to get the values there. What was missing was the channel, not the intent — and by the end of the afternoon the channel existed, built by the runtime itself, from one sentence.

Asked to bind, it also proposed a vocabulary for the part of a design system that does not exist yet: the scene-level roles a 2D system has never needed to name. It did that twice, independently, converging. That is a product behaviour worth knowing about, and the product currently tells nobody it is there.

Elsewhere the output is genuinely good: semantic, readable CSS instead of utility soup; labelled inputs with correct autocomplete and required attributes, so the accessibility floor on forms holds; a real project with declared dependencies that builds cleanly with its own toolchain; and compositional taste that does not read as generated.

The files used

Everything the runs consumed is public and served directly, so any arm in the appendix can be repeated against the same input rather than an approximation of it. The system is open-source under MIT.

FileUsed inURL
Full specification the 103 KB arms and the replication amaca.design/DESIGN.md
Compact instruction file the compact arm and all four binding arms amaca.design/downloads/AI-INSTRUCTIONS.md
Token source the reference the checker measures against amaca.design/styles/tokens.css
Tokens, DTCG — amaca.design/downloads/tokens.dtcg.json
Agent rules — amaca.design/downloads/AGENTS.md
Index for AI tools — amaca.design/llms.txt

The system itself, with the component gallery and the token tables: amaca.design. Source: github.com/angelomacaione/amaca-design.

The two files in the first two rows are the same design system in two shapes. That distinction is not cosmetic here — it is the subject of the parser observation above.

Appendix — the experiments, verbatim

Every arm below can be re-run. The canonical prompt was frozen at the first run and never edited; each arm differs from its neighbour by exactly one thing, and that thing is printed here in full.

The canonical prompt

Identical in every arm that generates a page. Not edited once, across two sessions a week apart.

Build a landing hero for a design-tools startup.

Layout: two columns on desktop, stacked on mobile. Full-viewport section.
Left column: a small status badge, a headline, one supporting paragraph,
an email capture field with a labelled input, and one primary call-to-action button.
Right column: an interactive 3D object that responds to cursor movement,
with a slow idle motion of its own.

Dark interface.

It is written that way on purpose. It describes a screen, not an aesthetic — no colours, no fonts, no measurements — so that anything visual in the output came from the specification rather than from the prompt. The labelled email field is deliberate: it exercises form components and the accessibility floor, which a headline-and-button hero would not. “Dark interface” stays because the system under test is dark-first; without it the no-specification baseline could come back light, and the comparison would be measuring two things at once.

The prompt that came before the method

Reported for completeness. No design system, no protocol, two iterations, not measured.

a landing page hero containing a 3d brain made of tiny opalescent cubes

And the second attempt, with the specification loaded, thirteen iterations, also not measured:

a landing page hero section containing a 3d brain slowly rotating made of tiny
opalescent cubes like a repeater, while hovering the cubes scale

And the third, same specification, new chat with no accumulated context, two iterations:

design a 3d opalescent 3 brain made with cubes, interactive on hover

The arms

ArmWhat was loadedAdded to the promptHeadline result
Free attemptnothingnothing — one line, no methodnot measured — see “Where this started”
Free attempt, with specthe design systemnothing — free prompt, 13 roundsnot measured — renders, does not steer
Free attempt, clean sessionthe design systemnothing — free prompt, 2 roundsnot measured — best of the three
Baselinenothingnothingnaming ~25%, no system anywhere
Full specificationthe 103 KB specificationnothingnaming 100%, 3D binding = mirror
Replicationthe same, one week laternothingindistinguishable from the above
Compact filea 5.6 KB instruction filenothingnaming 100%, values 100%
Plain-language requestthe design system (structured)a five-line request, belowmirror, but reasoned
Intent onlythe compact instruction fileone sentence, belowbinding goes live
Mechanism namedthe compact instruction fileone sentence, belowlive, naming drops to 74%
Three clausesthe compact instruction filethree clauses, belowlive, naming 100%, zero copies
Plain-language request

Asking for it in prose

Appended to the canonical prompt. It names the tokens but says nothing about how to reach them.

The 3D layer obeys the design system: the object's base material color is --obsidian-800,
its emissive accent is --magenta-500 kept under 5% of the visible surface,
the key light is neutral and the rim light is --magenta-500.
Idle rotation and cursor response use the system's duration and easing tokens.
Do not introduce colors, curves or durations that are not in the system.

Result: still a mirror — but a considered one. Five colours, all six durations converted to seconds, all four easing curves as exact control points, and the system's motion rules applied unprompted: a linear clock for the continuous loop because the specification says continuous loops are not eased, and a no-overshoot curve on the emissive because emissive is an effect property. It understood the grammar. It still could not read a value at runtime.

Intent only

One sentence

The whole difference between a copy and a binding.

The 3D layer must stay bound to the design system at runtime: never copy a token
value into JavaScript. If a token changes in the stylesheet, the 3D scene changes with it.

Result: a live resolver reading the page's custom properties, a mutation observer plus a colour-scheme listener that rebuild the scene when the stylesheet changes, a correct cubic-bézier solver that parses the curve out of the CSS string rather than retyping it as numbers, and a radius token driving the roundness of the geometry. None of that was requested by name. The mechanism was never mentioned in the prompt.

Mechanism named

Saying how

The 3D layer resolves its colours, durations and easings from the page's CSS custom
properties at runtime via getComputedStyle. No token value is hardcoded in JavaScript.

Result: live, and the hardcoded fallbacks nearly disappear — from eleven down to one. But naming fidelity falls to 74.2%: forced to be explicit about what it reads, the generator built two alias layers of its own, a semantic one and a scene-level one, both defined in terms of the system's real primitives but under names the system does not ship.

Three clauses

The version that holds

The 3D layer resolves every colour, duration and easing from the page's CSS custom
properties at runtime via getComputedStyle. No token value is hardcoded in JavaScript,
not even as a fallback: if a token is missing, fail loudly.

Read the design system's own custom-property names verbatim — --accent, --card, --border,
--bg, --text-secondary, --r-lg, --d-scene, --ease-decel. Do not create aliases or a
parallel namespace.

Honour prefers-reduced-motion inside the WebGL layer, not only in CSS.

Result: naming fidelity back to 100%, binding still live, zero hand-copied values anywhere including the fallbacks, a typed error thrown when a token is missing instead of a substituted literal, and reduced motion handled inside the scene with a visible state readout. Every clause did exactly one job, and no clause undid another.

Earlier session

The thirty-line contract

This is what it took the first time, before the shorter forms above were tried. Kept here because the mapping itself may be useful: none of these primitives were invented for the 3D layer — each is derived from something the 2D system already had.

3D token contract — derive the 3D layer from the design system, do not invent a parallel one:

MATERIAL   base color = --obsidian-800 · roughness 0.6 · metalness 0.1
           the surface scale (--obsidian-900 → -700) is the material's depth range
EMISSIVE   --sh-glow already encodes the brand glow → emissive = --magenta-500,
           intensity 0.25, and it obeys 85/10/5: emissive surface stays under 5%
DEPTH      the shadow scale is a Z scale: --sh-1 = 1 unit, --sh-2 = 4, --sh-3 = 12, --sh-4 = 24
LIGHT      85/10/5 is a light budget too: key light neutral (85%), fill from --tertiary-500 (10%),
           rim light --magenta-500 (5%). Never a white rim.
MOTION     idle and response ride --d-* / --ease-*, no other durations or curves.
           Spatial-vs-Effect holds in 3D exactly as in CSS: position/rotation/scale may use
           --ease-spring; material color, emissive and opacity may not.
SCALE      1 scene unit = 16px = --s-4, so 3D depth and 2D spacing share one grid.

Report at the end which of these you could honour and which you could not.

The last line was deliberate: ask the model to declare its own limits. It reported honouring six of the seven sections against a scene file that had not changed by a single byte. The round that converted the report into a real refactor added one instruction — paste back the actual lines, not a summary.

One constraint in that contract turned out not to be projectable at all: depth. The shadow tokens encode blur and spread, not a scalar distance, so there is no CSS property from which a Z value can be read. The model spotted that on its own and said so — correctly, and it is not an obvious observation. It then produced the appearance of compliance rather than refusing: an expression in which the token appears on both sides of a subtraction and cancels out, leaving the hand-tuned number it had before. A checker does not see that class of false compliance. A reader does.

Reproducing this

Load the specification, paste the canonical prompt into a fresh project, export the result, and measure the code — not the screenshot. The arms above differ by one variable each; anything else changed at the same time makes the comparison say nothing. Cost, for reference: a complete arm ran under a hundred credits.

Perimeter and rigour

Tested: generation, fidelity to a loaded specification, the 3D layer, and the structure of the exported project. Not tested: the editor, the canvas, the component library, collaboration, or mobile and XR export. Every claim above is bounded by that perimeter.

One run per condition, uncontrolled, temperature and seed unknown; one arm was replicated a week apart and is noted as such. The floor and ceiling are single measurements, not distributions. These are observations of what happened on a specific afternoon with a specific system — this is what I got when I did this, not Omma does X. Every number here is reproducible from the exported code with the same script.