VNHAX
vnhax
JuliusBrussee/cavemanPublic
29
507
3 branches · 14 tags
J
JuliusBrusseeperf: optimize Go stream lexer chunk parsing to sub-100µs latency
5b11c094 hours ago
cmd/caveman
cli: add start and ab-test terminal commands
4 hours ago
pkg/lexer
perf: streaming markdown lexer preserving code fences and diffs
4 hours ago
pkg/proxy
proxy: non-blocking reverse proxy for OpenAI/Anthropic APIs
yesterday
go.mod
chore: update Go runtime to 1.23
2 days ago
LICENSE
docs: MIT license
2 weeks ago
README.md
docs: add A/B benchmark graphs showing 40% token savings
3 hours ago
README.md
Verified Technical Teardown
🪨 #1 ON GITHUB TRENDING#1 ON HACKER NEWS (904 PTS)33.2% FEWER INPUT TOKENSCITED BY ADOBE RESEARCHAPACHE-2.0 LICENSE
Caveman - why many token when few do trick
JuliusBrussee / caveman — why many token when few do trick

Caveman: Token Optimization Proxy & Terse Voice Harness

Caveman make your AI agent say less and read less. Code stay exact. Brain still big.
The skill shrinks what the agent says by stripping conversational filler. The local proxy shrinks what it reads (logs, test runs, JSON dumps, web pages) by up to 99%.

33.2% fewer
Input tokens through proxy
54 Claude Code runs, 18/18 right
129.8× smaller
Web pages for the agent
caveman browse vs Playwright
1.4× to 2.4×
Cheaper execution cost
Adobe Research (8 frontier models)

Same Answer. 63 Tokens Become 20. Brain Still Big.

Normal Agent63 tokens
“The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.”
🪨 Caveman Agent20 tokens (-68%)
“New object ref each render, so React re-renders. Wrap the prop in useMemo.”
PICK YOUR CLUB: THREE TERSE INTENSITY MODES
/caveman20 tokens

“New object ref each render, so React re-renders. Wrap the prop in useMemo.”

/ultracave14 tokens

“Inline object prop, new ref, re-render. useMemo.”

/megacave13 tokens

“New ref triggers re-render. useMemo.”

THE VOICE

How Caveman Talks (Rules of Grammar & Silence)

Caveman is a disciplined voice, not broken grammar. Every reply follows strict linguistic rules:

RuleWhat It Means
Answer first[thing] [action] [reason]. [next step]. No greeting, no "let me", no recap, no "hope this helps".
One idea per sentenceBuilt on ASD-STE100 (controlled English for aircraft maintenance): 20 words max, active voice, one term per thing.
Meaning never droppedArticles can go. "not", "never", "no", "only" never go. Numbers and units stay exact.
Payload verbatimCode snippets, shell commands, file paths, and error messages are untouched character-for-character.
Quiet tool runsNo conversational chatter between tool calls. One line per phase, one line with the result.
Knows when to stopSecurity warnings, irreversible actions, multi-step ordering, and confused users get full sentences. Then grunt resumes.
Never performsNo cartoonish "me think", no caveman prefix. If caveman phrasing is not shorter, plain English wins.
Your prompts stay yoursUser prompts are never rewritten. Research shows compressing user prompts backfires and hurts answer accuracy.
Setup & Harness Distribution

Install (Works With 30+ Agents)

npx skills add JuliusBrussee/caveman -g

Works immediately in Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, Copilot, and 30+ more. Type /caveman to start. Say stop caveman to exit.

Harness-Specific Install Commands:
Claude Code Plugin:
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
Gemini CLI:
gemini extensions install https://github.com/JuliusBrussee/caveman
All agents on machine at once (macOS / Linux):
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.sh | bash
Windows (PowerShell 5.1+):
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.ps1 | iex

The Numbers (Empirical Benchmarks from Independent Labs)

Outside research labs first, then ours. Nothing rounded up:

WhoSetupResult
Adobe Research (CAVEWOMAN paper)Caveman-style output, 8 frontier models, 5 datasetsCost cut 1.4× to 2.4× per model, up to 3×
Elastic (Elasticsearch Labs)Caveman mode for Elasticsearch, 8 live MCP scenarios63.6% fewer response tokens (“Zero information loss”)
JetBrains Research86 real coding tasks, paired A/B evaluationNo measurable quality loss (p = 0.82), 8.5% fewer output tokens

The Proxy: 33.2% Fewer Input Tokens Across Whole Agent Sessions

File TypeFile Through CavemanFile SavedWhole Session (3 runs)Session Saved
CSV28,041 → 31498.9%165,823 → 74,48455.1%
Logs22,810 → 34898.5%148,807 → 74,06850.2%
YAML20,447 → 17899.1%132,124 → 71,02746.2%
Test Output18,806 → 20398.9%150,377 → 108,51427.8%
JSON18,837 → 28198.5%147,975 → 108,93926.4%
All 6 Types130,611 → 22,99482.4%885,793 → 591,67333.2%
* 18 of 18 test answers right across all tasks.

Big Rock: The Proxy

The skill shrinks what the agent says. The proxy shrinks what it reads: logs, test output, JSON dumps, git diffs, web pages. It runs locally on your machine with your own credentials:

npm install -g @caveman-ai/cli && caveman setup --install caveman claude        # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · pi
caveman learn

Ranks where tokens go from agent history on disk

caveman shrink -- pnpm test

Compresses noisy command outputs on the fly

caveman browse <url>

121 tokens instead of a 15,000-token web dump

caveman stats

Real token usage & savings audit across sessions

What You Get: Commands, Subagents & Work Patterns

Command / ComponentWhat It Does
/caveman · /ultracave · /megacaveThe voice, the grunt, the classical concise mode. /caveman status shows mode, /caveman off stops it.
/caveman-commitGenerates ultra-terse, high-signal one-line Conventional Commit messages.
/caveman-reviewOne finding per line: "L42: 🔴 null deref. Guard it." Zero prose filler.
/caveman-compress <file>Shrinks memory files (e.g. CLAUDE.md) by 46% and backs up the original.
/caveman-statsReads session logs and computes real token metrics for this coding session.
cavecrewSpecialized subagents that find, edit, and review code, then report back in caveman voice.
investigate-first · surgical-patch · safe-refactor · verify-and-stopBuilt-in work patterns that write less code and minimize context turnover.

The Skill, Unpacked (Output Compression)

The skill acts on the model's output generation loop. It suppresses redundant greetings, explanations of obvious code, and conversational sign-offs. It never compresses user inputs or prompts. Code blocks, terminal commands, identifiers, and file paths are emitted 100% verbatim.

The Proxy, Unpacked (Context & Input Compression)

The local proxy acts on model input context. When agents ingest 20,000-token JSON dumps or test outputs, Caveman stores the exact original in Caveman Context Recovery (CCR) on disk, provides a compact representation to the model, and passes through a retrieval handle if the agent needs the raw bytes.

How to Wrap an Agent & In Your Own Code

You can wrap an installed coding agent directly, or import Caveman middleware into your own application code:

Wrap Any Coding Agent
# Wrap Claude Code caveman claude # Wrap Codex, Gemini, Aider, OpenClaw caveman codex caveman gemini caveman openclaw

Rewrites provider loopback endpoint for the child process without modifying the agent core loop.

In Your Own Code (SDK & Middleware)
# TypeScript / Node.js npm install @caveman-ai/middleware @caveman-ai/sdk # Python (LangChain, OpenAI, LiteLLM) pip install 'caveman-middleware[langchain]' caveman-sdk

Drop-in middleware for Vercel AI SDK, LangChain, CrewAI, Pydantic AI, and raw HTTP calls.

HONEST NUMBERS

When NOT to Use Caveman (Net-Negative Cases)

Caveman saves tokens sometimes, but can cost tokens if misapplied. Turn it off if your workload hits these conditions:

  • Terse Coding Q&A: If asking 1-sentence questions, injecting the 1,000-token Caveman rule sheet into context will cost more tokens than the short answer saves.
  • Per-Request Pricing: If using models billed per request rather than per token (e.g., GitHub Copilot premium requests), shorter answers do not reduce billed cost.
  • Over-Aggressive Context Re-Injection: If tool-side harnesses repeatedly re-inject system prompts on retries, input token costs can overwhelm output savings.
  • Rule of Thumb: Always run an A/B test with caveman trial. If Caveman increases billed costs on your specific tasks, turn it off.
🔒 Privacy

The skill runs 100% locally on your machine and sends nothing. The CLI collects anonymous aggregate counts by default (never your code, prompts, or paths). Disable permanently via caveman telemetry off or DO_NOT_TRACK=1.

⚖️ License

Apache-2.0 License across the entire repository. Read it, fork it, ship it, host it. Free like mammoth on open plain.

📚 Academic Citation
@software{brussee2026caveman,
  author = {Brussee, Julius},
  title  = {Caveman},
  year   = {2026},
  url    = {https://github.com/JuliusBrussee/caveman}
}