Four reproducible observations
3 of 3
Agent file deletions do not persist
Three times, in three separate projects, a stray entry file at the project root broke the preview by loading an HTML file as a JavaScript module. Three times the agent diagnosed the cause correctly, in detail, and stated it had deleted the file. Three times the file was still there afterwards — once with the deletion claim on screen next to an export that still contained it, still carrying the exact line the message said had been removed.
Moving the entry files to the root worked immediately, both times it was tried. This reads as a scaffold behaviour, not a model failure: the reasoning is right and the delete is issued, and something downstream does not honour it. The cost lands on the user as rounds, credits, and the impression that the agent is not telling the truth.
8 and 9 rounds
The error message does not name the file
Unexpected token ‘<’ arrives without a path. The cause was an entry file loading an HTML document as a JavaScript module — the 3D scene was correct and running the whole time. The agents rewrote it anyway, repeatedly, because nothing told them where to look.
Eight rounds in one project, nine in another, to get past an error that was never in the generated code. Pointed at the right file by hand, it resolved in a single shot for sixteen credits. And that figure is a floor rather than a ceiling: the later rounds were not the runtime working unaided.
Naming the file in the error is probably the highest value-to-effort change in the whole debug loop, because every wrong attempt is spent from the user's credit balance. The user pays for the ambiguity.
4 occurrences
Reported work and the diff disagree
Beyond the deletions above, one round returned a six-section report describing a 3D refactor in specific technical detail — emissive intensity, light colours, easing sources — against a scene file that was byte-identical to the previous version.
The significant part is that the reasoning in that report was correct, and better than correct: the observation that shadow tokens encode blur and spread rather than a scalar distance, so there is no CSS property from which to read a depth, is true and not obvious. The model understood the contract. Then it reported having applied it. Asking for verbatim evidence — paste back the actual lines, not a summary — was what converted the next round from a report into a real refactor.
5 of 5
Reduced motion never reaches the WebGL layer
Across every run where it was not explicitly requested, there was no prefers-reduced-motion handling in the scene — with objects in perpetual rotation, orbiting rings and satellites. The CSS-level kill switch is emitted correctly and cannot reach WebGL. For a viewer with vestibular sensitivity, that hero cannot be turned off.
Nobody in this category appears to be handling it. Whoever does it first has a differentiator that enterprise buyers understand without explanation.
silent
The design-system parser accepts prose and invents tokens from it
The same system was loaded twice, in two forms: once as a structured document with a machine-readable header, once as the compact prose instruction file. The style panel shows what the platform made of each.
From the structured document, the token model is correct: the brand and neutral colours by name, the full neutral ramp, and a spacing scale reading s-1 4px, s-2 8px, s-3 12px, s-4 16px and up.
From the prose, the extracted spacing scale reads:
utline 2px
offset 3px
labels 10px
adding 14px
These are not tokens. They are fragments of words. utline is the tail of outline; adding is the tail of padding; offset comes from outline-offset. Each carries a number scraped from the sentence it appeared in. The colour panel shows an unnamed row of dark swatches.
The mechanism is visible in the output itself: this is substring matching applied to prose. outline-offset: 2px yields a token called utline worth 2px. padding: 14px yields adding. A pattern written for a structured file, run over sentences, returns the tail of every word it half-matches.
The generation is unaffected, and it is worth showing why rather than asserting it. The compact file states in its own first lines that the token values are inlined so that it works with no other files — the text contains --obsidian-950 #07090B and the rest in full. The language model reads that text. The parser is a second, separate consumer of the same file, and only the second one fails.
The generated code carries the proof. In the compact arm: sixty-six custom properties declared, sixty-six of them carrying the system's own names, zero invented, zero off-scale durations, two raw pixel values in the whole file. Had the generator been working from the parsed model, --utline and --adding would appear in the output and a 14px step would appear in the spacing. Neither does.
No error is raised and no warning is shown. The panel presents a broken system exactly as it presents a correct one, and a user pasting text instructions has no way to know which of the two they are looking at. That is what makes this quiet: the output looks right while the platform's own model of the system is nonsense, so nothing in the experience ever surfaces the gap.
across sessions
One three.js clock bug recurs
The animation loop calls the elapsed-time accessor before the delta accessor, which zeroes the delta — measured at 0.000024 radians instead of 0.302 after two and a half seconds. Anything driven by delta barely moves. It appeared in the first session, was corrected, reappeared in the next project, and appeared again a week later in a fresh project. It is not that it does not learn between projects: it does not learn between sessions.