Caveman
Also known as: caveman proxy, token-crushing proxy
Caveman sits between your coding agent and the model API (the interface your tool uses to send requests to a language model). Before a prompt goes out, Caveman rewrites it with heavy abbreviation: stripping vowels, collapsing common phrases, and dropping connective text the model can reconstruct from context. The output is often illegible to a human but readable to a sufficiently capable LLM.
The motivation is cost and latency. At scale, even a modest reduction in prompt length across thousands of agent turns adds up to meaningful savings on inference bills. Caveman approaches this from the opposite direction to Ponytail: where Ponytail reduces how much code the agent generates, Caveman reduces how much text the agent sends. They are often used together.
The project reached 109,000 GitHub stars by early October 2026, making it one of the most-starred agent-tooling repos of the year. The approach has tradeoffs: extremely compressed prompts can drift in meaning, and debugging a caveman-proxied session requires an extra step to reconstruct what was actually sent. It is best suited to high-volume, well-tested workflows where cost pressure outweighs the overhead of that debugging layer.