
Compression middleware that improves LLM outputs.

Official pages and third-party coverage in one index.
thetokencompany.com11 items across 10 mapped pages · Crawled Sep 24, 2026
Official pages and third-party coverage in one index
Bear-2-Safety compresses input to safety classifiers by up to 30% while preserving or improving F1 across Llama-Guard, ShieldGemma, and Gemini. Unsafe content is preserved at 95-100% retention.
Bear-2 compression improved CoQA accuracy from 93.3% to 95.3% while cutting tokens by 8.2%. Removing filler helps the model focus.
Pax Historia, processing 193B tokens/month on OpenRouter, ran a 268K-vote model arena with bear-1.1 compression. Compressed models scored higher and A/B tests showed +5% purchase amount lift.
Simple pricing for LLM input compression — you only pay for the tokens we remove from your input.
Talk to The Token Company about cutting LLM API costs, deploying prompt compression in production, or enterprise pricing for GPT, Claude, and Gemini workloads.
Cut OpenAI, Anthropic, and Gemini API costs with accuracy held flat — or take a smaller cut and lift accuracy by several points instead. The bear-2 prompt compression API strips low-signal tokens from your inputs before they hit the LLM. Works with GPT, Claude, Gemini, and any chat completion endpoint.
How The Token Company collects, uses, discloses, and otherwise processes personal data.
Helonic runs AI agents on construction drawings with long prompts. bear-1.2 compression trims ~47K tokens per prompt while preserving every critical detail.
The Token Company API Terms and Conditions governing access to and use of the TTC API.
Cut OpenAI, Anthropic, and Gemini API costs with accuracy held flat — or take a smaller cut and lift accuracy by several points instead. The bear-2 prompt compression API strips low-signal tokens from your inputs before they hit the LLM. Works with GPT, Claude, Gemini, and any chat completion endpoint.