Benchmark

AK Render token-efficiency benchmark

70–73% less presentation-context guidance in our AK Render benchmark. Presentation guidance dropped from about 23.7k–26.6k tokens to about 6.6k–7.1k on five representative HTML-producing tasks.

Measures presentation-context guidance, not total model token usage.

@bestagentkits/render@0.2.0 · measured October 6, 2026 · verified October 6, 2026

Results by task

Estimated presentation-context tokens per task
TaskLegacy guidance (tokens)AK Render (tokens)Reduction
Explain a system as an HTML page23,6986,95870.6%
Brainstorm options as an HTML page23,6986,62572%
Plan review page (Engineer)23,6987,13169.9%
Plan review page (Marketing)23,6987,13169.9%
Diff preview page26,5537,14973.1%

Why it works

  1. 01

    Describe intent

    The agent writes a short declarative Page Spec instead of regenerating HTML, CSS and JavaScript.

  2. 02

    Load only what is needed

    A compact block catalog plus descriptions of only the blocks in use, instead of full schema injection.

  3. 03

    Keep artifacts out of context

    The compiler writes the HTML to disk; the agent gets back a short summary and diagnostics.

Method

  • Tokens are estimated from character counts at 4 characters per token. No tokenizer or live model run is involved.
  • Before: the hand-written HTML guidance from AgentKit’s ak-preview references (shared contract, design guidelines, CSS patterns, libraries, responsive navigation). The diff-preview task follows that reference-loading table directly and also loads three ak-preview routing references; the explain, brainstorm and plan tasks use the ak-preview document references as a proxy for their own hand-written HTML route.
  • After: the same shared contract plus the AK Render skill, its block catalog, and the block descriptions for the blocks the matching fixture uses.
  • Both sides read the same AgentKit revision. The compiled page is written to disk and never re-enters model context.
  • Compiler @bestagentkits/[email protected], measured October 6, 2026. Artifact published in ak-render commit 91cfd56 (generated at f5b680710); AgentKit revision 9e322f928.

Limitations

  • Measures presentation-context guidance only, not total model token usage.
  • Excludes the invoking skill body, task content, model output tokens, tool-call tokens, retries, and wall-clock time.
  • Covers five representative HTML-producing tasks. Other AgentKit workflows are not claimed to save the same amount.
  • AK Render ships as its own agent skill and plugin. AgentKit’s HTML-producing skills still write HTML by hand by default.

Next benchmark

  • Live A/B runs in Claude Code and Codex on the same tasks.
  • Record input, cached, output, and tool-call tokens per run.
  • Count retries and repair loops, and record wall-clock time.
  • Judge output quality and completeness blind against the hand-written route.

Sources