What it actually does
The premise is a joke that turns out to be a real optimisation. Your agent's prose is padding: transitions, hedges, pleasantries. The information density is low and you pay for every token.
Caveman strips it. The agent drops filler and answers in tight caveman-speak — 65% fewer output tokens on prose, 8.5% on long-horizon agentic coding runs. Crucially, code, commands, and error messages stay byte-for-byte exact. It compresses the talking, not the artifacts.
It installs across Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, and thirty-plus other agents, with configurable levels if full caveman is too much.
The distinction from a skill like Ponytail is worth being clear about: Ponytail reduces how much code gets written, Caveman reduces how much prose surrounds it. They address different lines on the same bill and can be used together. MIT, 97,000 stars.
Who it's for
- 01
Anyone whose output token spend is a real line item
- 02
Developers who want the answer without the essay
- 03
Teams running high-volume agent workloads where percentages compound
- 04
People who find agent politeness actively slows them down
Where it earns its keep
- Cutting output token cost across a team with a one-time install
- Making agent responses scannable during focused work
- Reducing spend on long agentic runs where prose accumulates
- Keeping code output exact while compressing everything around it
- Tuning verbosity per context using the level settings
Use it, or skip it
Reach for it when
- Output tokens are a meaningful cost
- You want answers, not explanations, in day-to-day work
- Your agent workloads are high volume so small percentages matter
- You can accept unusual phrasing in exchange for density
Skip it when
- Output is customer-facing — caveman-speak in a support reply is not the impression you want
- You are learning and the explanation is the point
- Your team would find it genuinely irritating rather than funny
- You need agent output to be copied directly into professional documents
10 automations
Ideas, not tutorials. Each one is work a team does by hand today.
- 01Finance
Fleet-wide token reduction
Install across every agent in the team and report the aggregate monthly saving to finance.
- 02Operations
Internal-only verbosity profile
Apply it to internal tooling while leaving customer-facing generation untouched.
- 03Engineering
Long-run cost control
Enable specifically for long autonomous runs where prose accumulates across hundreds of turns.
- 04Engineering
Level tuning by context
Use a light level for design discussion and a heavy one for repetitive execution work.
- 05Finance
Stacked with code reduction
Combine with a code-minimising skill and measure the compounded effect on total spend.
- 06Operations
Status update compression
Apply to automated status posts so channels stay readable and cheap to generate.
- 07Engineering
CI output slimming
Reduce the prose in automated CI commentary while keeping error output verbatim.
- 08Engineering
Benchmark validation
Measure the claimed 65% against your own workload rather than assuming it transfers.
- 09Founders
Mobile-friendly output
Shorter responses are far easier to read when checking agent work from a phone.
- 10Engineering
Rate-limit relief
Fewer output tokens per turn means fewer rate-limit stalls on constrained plans.
Want one of these running by Friday?
LimeDock builds these as real workflows inside your stack — deployed to your cloud, wired into your Slack and CRM, with the code in your repo. You pay a build fee and your own API keys, nothing else.
Source
Repository stats were read from the GitHub API and reflect the last time we refreshed this entry. The editorial breakdown above is LimeDock’s own analysis — we are not affiliated with JuliusBrussee.
https://github.com/JuliusBrussee/caveman