Skip to content

TokenPak

A local proxy that packs LLM context before it reaches the API, with per-request records of what changed.

This page is for developers evaluating or getting started with TokenPak. TokenPak sits between your AI tools and the upstream LLM provider, with its proxy listening on 127.0.0.1. It deterministically packages context (Prompt Packing), routes requests, evaluates configured Spend Guard limits before provider send, and records request results locally. Provider-bound requests still travel to the selected upstream provider; TokenPak operates no cloud relay and requires no application code changes.

v1.19.0

The commands below, Quick Start, extended API reference, and Docker guide describe TokenPak v1.19.0, the currently published release on PyPI (pip install tokenpak). The separate Installation page retains older-release guidance; use the Quick Start for the current setup path. Other pages with explicit version pins describe the release line named on that page.


What ships in the OSS beta

  • Prompt Packing pipeline — deterministic context reduction on real agent workloads; reduction pinned to an agent-style CI fixture (reproduce with make benchmark-headline); provider-cached flows show lower incremental gains. Measure your own with tokenpak savings.
  • Local proxy on 127.0.0.1 — processing and records stay local; provider-bound prompts and credentials are sent to the upstream provider you configure, not to a TokenPak cloud service.
  • Spend Guard — pre-send circuit breaker with rolling caps; blocks runaway requests before they reach the provider and returns a clear release directive.
  • Nine client integrations — Claude Code, Cursor, Cline, Continue, Aider, Codex CLI, OpenAI SDK, Anthropic SDK, LiteLLM.
  • Savings Ledger + local dashboard — every request logged to a local SQLite store with causal attribution; TUI + web dashboard.
  • Vault indexing + semantic search — index your codebase, search without an LLM call.
  • TIP-1.0 protocol contracts — canonical headers, metadata fields, capability labels, manifest schemas. Conformance gate runnable via tokenpak doctor --conformance.
  • Pak recall (read-only) — storage, FTS, tokenpak pak inspect. Scoring and assembly are not part of the OSS beta.
  • Three built-in setup profiles and 50+ compression recipes — minimal, balanced, and aggressive profiles plus customizable packaged YAML recipes.

Quick start

pip install tokenpak
tokenpak setup --start
# Then point your client at http://127.0.0.1:8766

5-minute Quick StartOlder-release installation guide


Documentation map

Section What it covers
Installation Older-release installation guidance; use the Quick Start for v1.19.0
Quick Start Setup wizard, client integration, first savings in 5 minutes
Configuration How configuration works (env vars + YAML, precedence)
Environment Variables Complete TOKENPAK_* reference
CLI Reference Every verb, flag, and exit code (auto-generated)
Architecture Three planes, modular subsystems, proxy-centered design
Savings How TokenPak attributes savings causally
Security Auth tokens, TLS, audit logging, data privacy
Troubleshooting Common symptoms and fixes that work
Known Issues Current OSS-beta limitations
FAQ General questions
Recall overview Paks, reason codes, risk flags — the OSS data plane
Client Guides Per-client integration walkthroughs (Claude Code, Cursor, Cline, Continue, Aider, Codex CLI, Gemini CLI, OpenAI/Anthropic SDK)

Source and package