Benchmark
AK Render token-efficiency benchmark
70–73% less presentation-context guidance in our AK Render benchmark. Presentation guidance dropped from about 23.7k–26.6k tokens to about 6.6k–7.1k on five representative HTML-producing tasks.
Measures presentation-context guidance, not total model token usage.
@bestagentkits/render@0.2.0 · measured October 6, 2026 · verified October 6, 2026
Results by task
| Task | Legacy guidance (tokens) | AK Render (tokens) | Reduction |
|---|---|---|---|
| Explain a system as an HTML page | 23,698 | 6,958 | 70.6% |
| Brainstorm options as an HTML page | 23,698 | 6,625 | 72% |
| Plan review page (Engineer) | 23,698 | 7,131 | 69.9% |
| Plan review page (Marketing) | 23,698 | 7,131 | 69.9% |
| Diff preview page | 26,553 | 7,149 | 73.1% |
Why it works
01
Describe intent
The agent writes a short declarative Page Spec instead of regenerating HTML, CSS and JavaScript.
02
Load only what is needed
A compact block catalog plus descriptions of only the blocks in use, instead of full schema injection.
03
Keep artifacts out of context
The compiler writes the HTML to disk; the agent gets back a short summary and diagnostics.
Method
- Tokens are estimated from character counts at 4 characters per token. No tokenizer or live model run is involved.
- Before: the hand-written HTML guidance from AgentKit’s ak-preview references (shared contract, design guidelines, CSS patterns, libraries, responsive navigation). The diff-preview task follows that reference-loading table directly and also loads three ak-preview routing references; the explain, brainstorm and plan tasks use the ak-preview document references as a proxy for their own hand-written HTML route.
- After: the same shared contract plus the AK Render skill, its block catalog, and the block descriptions for the blocks the matching fixture uses.
- Both sides read the same AgentKit revision. The compiled page is written to disk and never re-enters model context.
- Compiler @bestagentkits/[email protected], measured October 6, 2026. Artifact published in ak-render commit 91cfd56 (generated at f5b680710); AgentKit revision 9e322f928.
Limitations
- Measures presentation-context guidance only, not total model token usage.
- Excludes the invoking skill body, task content, model output tokens, tool-call tokens, retries, and wall-clock time.
- Covers five representative HTML-producing tasks. Other AgentKit workflows are not claimed to save the same amount.
- AK Render ships as its own agent skill and plugin. AgentKit’s HTML-producing skills still write HTML by hand by default.
Next benchmark
- Live A/B runs in Claude Code and Codex on the same tasks.
- Record input, cached, output, and tool-call tokens per run.
- Count retries and repair loops, and record wall-clock time.
- Judge output quality and completeness blind against the hand-written route.