blog·Thu, Jul 2, 04:02 PM·Confidence 90%high qualitypublic
Prompt Compression API — Cut LLM Token CostsCut OpenAI, Anthropic, and Gemini API costs with accuracy held flat — or take a smaller cut and lift accuracy by several points instead. The bear-2 prompt compression API strips low-signal tokens from your inputs before they hit the LLM. Works with GPT, Claud…
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, Jun 29, 09:16 AM·Confidence 90%high qualitypublic
One of the biggest token consumers globally improved quality by removing context bloatPax Historia, processing 193B tokens/month on OpenRouter, ran a 268K-vote model arena with bear-1.1 compression. Compressed models scored higher and A/B tests showed +5% purchase amount lift.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, Jun 29, 09:16 AM·Confidence 90%high qualitypublic
How input compression enabled world class performance for long running agentsHelonic runs AI agents on construction drawings at near million-token prompts. bear-1.2 compression trims tokens while preserving every critical detail.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, Jun 29, 09:16 AM·Confidence 90%high qualitypublic
Compressing Conversational Context Without Losing the ThreadBear-2 compression improved CoQA accuracy from 93.3% to 95.3% while cutting tokens by 8.2%. Removing filler helps the model focus.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, Jun 29, 09:16 AM·Confidence 90%high qualitypublic
Introducing Bear-2-SafetyBear-2-Safety compresses input to safety classifiers by up to 30% while preserving or improving F1 across Llama-Guard, ShieldGemma, and Gemini. Unsafe content is preserved at 95-100% retention.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, May 11, 04:39 AM·Confidence 90%high qualitypublic
Why Your RAG App's Token Bill Is So HighMost RAG applications over-fetch context by 3-5x. Break down where your retrieval tokens go and how compression, reranking, and chunk hygiene cut costs.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, May 11, 04:39 AM·Confidence 90%high qualitypublic
Cut Your LLM API Costs by 70% Without Losing QualityA practical framework for reducing LLM API costs across system prompts, conversation history, RAG context, tool schemas, and output. Real pricing math included.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, May 11, 04:39 AM·Confidence 90%high qualitypublic
bear-1.1: Improved LLM Compression Model - The Token Companybear-1.1 is the latest LLM input compression model from The Token Company. An improved version of bear-1 with better accuracy preservation and faster compression speeds. Reduce AI costs by 3x.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗blog·Mon, May 11, 04:39 AM·Confidence 90%high qualitypublic
bear-1: LLM Input Compression Model - The Token Companybear-1 is an LLM input compression model by The Token Company that reduces tokens by 66% while maintaining or improving accuracy. Reduce AI costs by 3x with a single API call.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗