Caveman: Token Optimization Proxy & Terse Voice Harness
Caveman make your AI agent say less and read less. Code stay exact. Brain still big.
The skill shrinks what the agent says by stripping conversational filler. The local proxy shrinks what it reads (logs, test runs, JSON dumps, web pages) by up to 99%.
Same Answer. 63 Tokens Become 20. Brain Still Big.
“The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.”
“New object ref each render, so React re-renders. Wrap the prop in useMemo.”/caveman20 tokens“New object ref each render, so React re-renders. Wrap the prop in useMemo.”
/ultracave14 tokens“Inline object prop, new ref, re-render. useMemo.”
/megacave13 tokens“New ref triggers re-render. useMemo.”
How Caveman Talks (Rules of Grammar & Silence)
Caveman is a disciplined voice, not broken grammar. Every reply follows strict linguistic rules:
Install (Works With 30+ Agents)
npx skills add JuliusBrussee/caveman -g
Works immediately in Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, Copilot, and 30+ more. Type /caveman to start. Say stop caveman to exit.
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
gemini extensions install https://github.com/JuliusBrussee/caveman
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.sh | bash
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.ps1 | iex
The Numbers (Empirical Benchmarks from Independent Labs)
Outside research labs first, then ours. Nothing rounded up:
The Proxy: 33.2% Fewer Input Tokens Across Whole Agent Sessions
| File Type | File Through Caveman | File Saved | Whole Session (3 runs) | Session Saved |
|---|---|---|---|---|
| CSV | 28,041 → 314 | 98.9% | 165,823 → 74,484 | 55.1% |
| Logs | 22,810 → 348 | 98.5% | 148,807 → 74,068 | 50.2% |
| YAML | 20,447 → 178 | 99.1% | 132,124 → 71,027 | 46.2% |
| Test Output | 18,806 → 203 | 98.9% | 150,377 → 108,514 | 27.8% |
| JSON | 18,837 → 281 | 98.5% | 147,975 → 108,939 | 26.4% |
| All 6 Types | 130,611 → 22,994 | 82.4% | 885,793 → 591,673 | 33.2% |
Big Rock: The Proxy
The skill shrinks what the agent says. The proxy shrinks what it reads: logs, test output, JSON dumps, git diffs, web pages. It runs locally on your machine with your own credentials:
npm install -g @caveman-ai/cli && caveman setup --install caveman claude # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · pi
caveman learnRanks where tokens go from agent history on disk
caveman shrink -- pnpm testCompresses noisy command outputs on the fly
caveman browse <url>121 tokens instead of a 15,000-token web dump
caveman statsReal token usage & savings audit across sessions
What You Get: Commands, Subagents & Work Patterns
The Skill, Unpacked (Output Compression)
The skill acts on the model's output generation loop. It suppresses redundant greetings, explanations of obvious code, and conversational sign-offs. It never compresses user inputs or prompts. Code blocks, terminal commands, identifiers, and file paths are emitted 100% verbatim.
The Proxy, Unpacked (Context & Input Compression)
The local proxy acts on model input context. When agents ingest 20,000-token JSON dumps or test outputs, Caveman stores the exact original in Caveman Context Recovery (CCR) on disk, provides a compact representation to the model, and passes through a retrieval handle if the agent needs the raw bytes.
How to Wrap an Agent & In Your Own Code
You can wrap an installed coding agent directly, or import Caveman middleware into your own application code:
# Wrap Claude Code caveman claude # Wrap Codex, Gemini, Aider, OpenClaw caveman codex caveman gemini caveman openclaw
Rewrites provider loopback endpoint for the child process without modifying the agent core loop.
# TypeScript / Node.js npm install @caveman-ai/middleware @caveman-ai/sdk # Python (LangChain, OpenAI, LiteLLM) pip install 'caveman-middleware[langchain]' caveman-sdk
Drop-in middleware for Vercel AI SDK, LangChain, CrewAI, Pydantic AI, and raw HTTP calls.
When NOT to Use Caveman (Net-Negative Cases)
Caveman saves tokens sometimes, but can cost tokens if misapplied. Turn it off if your workload hits these conditions:
- Terse Coding Q&A: If asking 1-sentence questions, injecting the 1,000-token Caveman rule sheet into context will cost more tokens than the short answer saves.
- Per-Request Pricing: If using models billed per request rather than per token (e.g., GitHub Copilot premium requests), shorter answers do not reduce billed cost.
- Over-Aggressive Context Re-Injection: If tool-side harnesses repeatedly re-inject system prompts on retries, input token costs can overwhelm output savings.
- Rule of Thumb: Always run an A/B test with
caveman trial. If Caveman increases billed costs on your specific tasks, turn it off.
The skill runs 100% locally on your machine and sends nothing. The CLI collects anonymous aggregate counts by default (never your code, prompts, or paths). Disable permanently via caveman telemetry off or DO_NOT_TRACK=1.
Apache-2.0 License across the entire repository. Read it, fork it, ship it, host it. Free like mammoth on open plain.
@software{brussee2026caveman,
author = {Brussee, Julius},
title = {Caveman},
year = {2026},
url = {https://github.com/JuliusBrussee/caveman}
}