Letting a CLI handle layout so the model only writes content is a smart tradeoff
on: QingYunA/answer-me-with-html
The core insight here is simple: when an agent writes HTML by hand, most of what it produces is not content. Coordinates, paths, CSS, tag scaffolding — the README measured this across several hand-written pages and found that actual text accounted for only about a fifth of the tokens. The rest is structural boilerplate the model recomputes from scratch every time.
Answer me with HTML reroutes that work. The model writes a short Markdown draft — section headers, a few fenced code blocks with diagram types like sequence or flow, plain prose. A bundled CLI called am handles everything else: it places panels in a grid, runs dagre for flow-chart layout, spaces sequence diagrams by label width, picks a theme, and emits a single self-contained HTML file with no CDN dependencies. The result opens offline and embeds the Markdown source that produced it.
The token reduction is the headline number: across the benchmark questions, the model wrote roughly one-eighth as many tokens compared to generating HTML directly, and the wall-clock time dropped by about 2.6 times. The cost savings are smaller — around 15% — because every turn still reads the system prompt and conversation history regardless of how much the model writes. The README is honest about this asymmetry.
Explainer videos follow the same pattern. The model writes a draft with one narration line per beat; am video turns it into an animated player page with chapter navigation, speed control, and in-browser video export. The README cites a small test showing roughly 18 times fewer output tokens and 12 times faster generation compared to asking for video HTML directly. The voice plays from inside the page, so it works offline.
The writing check is an unexpected inclusion. Every render runs the draft against rules adapted from ASD-STE100, the controlled-English standard originally written for aircraft maintenance manuals. It flags long sentences, passive voice, and wordy phrases. By default it only warns; strict mode refuses drafts that fail.
Installation targets more than 70 agents through the npx skills installer, and the Claude Code plugin path is a two-command process. The only runtime dependency beyond Node.js 20 is marked for Markdown parsing and dagre for layout, both bundled into a single am.mjs file.
The architectural bet is that diagram layout and CSS are deterministic enough to be handled by a CLI, while content judgment belongs to the model. That division holds up well for the component types the skill supports — flow charts, sequence diagrams, timelines, ER diagrams, annotated text. It would break down for anything requiring visual creativity the model actually needs to express. But for technical explanation, which is most of what agents do, the tradeoff looks sound.
Offloads diagram layout and CSS to a CLI so the model writes only content — a clean division that makes agent-generated pages faster and cheaper without sacrificing quality.