← Back to glossary
+Suggest a term
Tool·Builder Tools·Added today

Caveman

Also known as: caveman proxy, token-crushing proxy

An open-source proxy layer for coding agents that compresses prompts using extreme abbreviation before they reach the model. The goal is to shrink token usage and cut inference costs, at the expense of human readability.

Caveman sits between your coding agent and the model API (the interface your tool uses to send requests to a language model). Before a prompt goes out, Caveman rewrites it with heavy abbreviation: stripping vowels, collapsing common phrases, and dropping connective text the model can reconstruct from context. The output is often illegible to a human but readable to a sufficiently capable LLM.

The motivation is cost and latency. At scale, even a modest reduction in prompt length across thousands of agent turns adds up to meaningful savings on inference bills. Caveman approaches this from the opposite direction to Ponytail: where Ponytail reduces how much code the agent generates, Caveman reduces how much text the agent sends. They are often used together.

The project reached 109,000 GitHub stars by early October 2026, making it one of the most-starred agent-tooling repos of the year. The approach has tradeoffs: extremely compressed prompts can drift in meaning, and debugging a caveman-proxied session requires an extra step to reconstruct what was actually sent. It is best suited to high-volume, well-tested workflows where cost pressure outweighs the overhead of that debugging layer.

This definition is AI-generated and refreshed weekly. It may contain inaccuracies. Use your own judgment, especially for production decisions.
Related terms
PonytailContext engineeringToken BudgetContext compactionAgentic coding