API Documentation
Compress prompts programmatically via our REST API or npm SDK.
Claude Code hook (retired)
The TokenShrink Claude Code hook is retired. Claude Code's UserPromptSubmit hooks can add context to a prompt, but they can't replace it. So the hook never reduced what Claude received, and the savings counter it showed was not real.
The hook installed by the one-line installer also sent each prompt you typed to tokenshrink.com for compression. TokenShrink's database stored only counts (word and token totals), never the prompt text. The installer and the hook download now do nothing.
If you installed the hook, please remove it:
- Open
~/.claude/settings.jsonand delete theUserPromptSubmitentry whose command containstokenshrink-compress. Also delete anyPreToolUseorSessionStartentry whose command containstokenshrink-. - Run:
rm -f ~/.claude/hooks/tokenshrink-*.js ~/.claude/.tokenshrink-saved ~/.claude/.tokenshrink-log.jsonl - If you set up the session vocabulary, also run:
rm -f ~/.claude/session-vocab.json - If you added a status line that reads
~/.claude/.tokenshrink-saved, remove it from~/.claude/settings.json.
Cursor Integration
Use TokenShrink as a preprocessor for Cursor's AI features. Add to your project:
# Install the SDK
npm install tokenshrink
# Use in your Cursor rules or project config
import { compress } from 'tokenshrink';
// Compress before sending to Cursor's AI
const { compressed, stats } = compress(yourPrompt);Quick start
1. Get your API key
Sign up (free), then generate an API key from your dashboard.
2. Compress via API
curl -X POST https://tokenshrink.com/api/compress \
-H "Content-Type: application/json" \
-H "x-api-key: ts_live_your_key_here" \
-d '{
"text": "Your long prompt text here...",
"domain": "auto"
}'3. Or use the SDK (v2.0)
npm install tokenshrink
import { compress } from 'tokenshrink';
// Compress a prompt — runs locally, no API call needed
const result = compress('Your long prompt...');
console.log(result.compressed);
console.log(result.stats.tokensSaved); // Real token savings
console.log(result.stats.originalTokens); // Original token count
console.log(result.stats.totalCompressedTokens); // Compressed token count
// Optional: plug in a real tokenizer for exact counts
import { encode } from 'gpt-tokenizer';
const result2 = compress('Your long prompt...', {
tokenizer: (text) => encode(text).length
});
// Use with any LLM provider
import OpenAI from 'openai';
const openai = new OpenAI();
const res = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'system', content: result.compressed }],
});Compression strategies
Phrase compression
Variable savingsRemoves filler words and replaces verbose phrases with concise alternatives. Runs automatically on every prompt — no configuration needed.
"I would like you to please make sure to check the file" → "check the file"
Domain compression
Experimental savingsAuto-detects your tech stack from package.json and applies domain-specific abbreviations. Supports React, Node.js, Python, Supabase, SQL, TypeScript, Docker, Tailwind.
"useCallback" → "UCB", "getServerSideProps" → "SSP"
Session vocabulary
Experimental savingsLocal integration experiment. At session start, builds a project-specific codebook. Both you and the AI use the same abbreviations. Cost is paid once, amortized across every message.
"/Volumes/AI-Models/" → "SSD/" on every prompt, forever
API Reference
Headers
Request body
{
"text": "string (required) — the text to compress",
"domain": "string (optional) — auto|code|medical|legal|business"
}Response
{
"compressed": "string — full compressed text with Rosetta header",
"rosetta": "string — just the decoder header",
"stats": {
"originalWords": 150,
"compressedWords": 42,
"rosettaWords": 18,
"totalCompressedWords": 60,
"originalTokens": 168,
"compressedTokens": 45,
"rosettaTokens": 22,
"totalCompressedTokens": 67,
"ratio": 2.5,
"tokensSaved": 101,
"dollarsSaved": 0.05,
"strategy": "domain",
"domain": "code",
"tokenizerUsed": "built-in"
}
}Returns your current usage stats, monthly history, and recent compressions. Requires authentication (session or API key).
Rate limits
Token counting
v2.0 uses real token counts instead of word estimates. By default, TokenShrink uses a precomputed lookup table based on cl100k_base (GPT-4). For exact counts with your specific model, pass a custom tokenizer:
import { compress } from 'tokenshrink';
import { encode } from 'gpt-tokenizer';
const result = compress(text, {
tokenizer: (text) => encode(text).length
});Compression domains
Set domain to optimize compression for specific content types. Default is auto.
autoAutomatically detects the best strategycodeProgramming and technical documentationmedicalMedical records, clinical noteslegalContracts, legal documentsbusinessBusiness communications, reports