Use TokenPak with Gemini CLI¶
Not verified by current release evidence. The first-class integrations are Claude Code and Codex, and the tested adapters are the OpenAI SDK, Anthropic SDK and LiteLLM; this guide describes an expected setup for Gemini CLI.
This guide is for developers using Google's Gemini CLI who want to try routing Gemini requests through TokenPak.
Gemini CLI supports a custom Gemini API base URL through GOOGLE_GEMINI_BASE_URL. Point that variable at TokenPak's local proxy.
What you need before starting:
- Gemini CLI (
@google/gemini-cli) installed - TokenPak installed locally
GEMINI_API_KEYfor Google AI Studio- A shell where you can export environment variables before launching
gemini
Copy-paste setup¶
pip install tokenpak
tokenpak setup
curl -s http://localhost:8766/health | python3 -m json.tool
Then launch Gemini CLI from the same shell:
export GEMINI_API_KEY="your-gemini-api-key"
export GOOGLE_GEMINI_BASE_URL="http://localhost:8766"
gemini -p "Reply with one sentence confirming Gemini CLI is routed through TokenPak."
Do not add /v1 to GOOGLE_GEMINI_BASE_URL. Gemini CLI sends Google Generative AI paths such as /v1beta/models/...:generateContent; TokenPak detects those paths and routes them to the Google adapter.
1. Start TokenPak¶
pip install tokenpak
tokenpak setup
tokenpak setup detects provider keys, creates ~/.tokenpak/config.yaml, and starts the proxy on port 8766. You should see:
TokenPak proxy listening on http://localhost:8766
Confirm the proxy is healthy:
curl -s http://localhost:8766/health | python3 -m json.tool
Expected response shape:
{
"status": "ok",
"uptime_seconds": 3,
"version": "1.25.1",
"requests_total": 0,
"requests_errors": 0,
"compression_ratio_avg": 0.0
}
If status is not "ok", run tokenpak status before launching Gemini CLI.
2. Launch Gemini CLI through TokenPak¶
Export the Gemini key and base URL in the same terminal where you run gemini:
export GEMINI_API_KEY="your-gemini-api-key"
export GOOGLE_GEMINI_BASE_URL="http://localhost:8766"
gemini -p "Say hello through TokenPak."
Use GOOGLE_GEMINI_BASE_URL for Google AI Studio / Gemini API traffic. If you intentionally use Vertex AI mode, Gemini CLI also supports GOOGLE_VERTEX_BASE_URL, but Vertex setup has separate project and location requirements; this guide focuses on the Google AI Studio path.
3. Verify traffic is routed through TokenPak¶
After Gemini CLI returns a response:
tokenpak status
You should see at least one recent request. You can also check /health again:
curl -s http://localhost:8766/health | python3 -m json.tool
If requests_total remains 0, Gemini CLI did not inherit GOOGLE_GEMINI_BASE_URL; see Troubleshooting.
4. Check your usage and savings¶
After a few Gemini CLI prompts:
tokenpak cost --week # spend by model
tokenpak savings # recorded token savings; zero is a valid result
The default proxy preserves conversation turns, so a forwarded request can truthfully report zero tokens saved. Explicit context tools can reduce eligible content; measure their effect with tokenpak savings.
Troubleshooting¶
requests_total stays 0 after Gemini CLI responds¶
Gemini CLI is not using the TokenPak base URL. Confirm:
GOOGLE_GEMINI_BASE_URLis exported in the same shell that runsgemini.- The value is exactly
http://localhost:8766. - You did not include
/v1or/v1betain the base URL. - You restarted Gemini CLI after changing the variable.
Run this in the same shell before launching Gemini CLI:
printf '%s\n' "$GOOGLE_GEMINI_BASE_URL"
Proxy not started — Gemini CLI shows connection errors¶
Verify the proxy directly:
curl -s http://localhost:8766/health
If this returns Connection refused, start the proxy:
tokenpak serve
You can also re-run tokenpak setup if this is your first install.
Port collision — proxy fails to start on 8766¶
If 8766 is already in use:
lsof -i :8766
Stop the conflicting process, then restart TokenPak. Alternatively, run TokenPak on another port and update Gemini CLI's base URL:
TOKENPAK_PORT=8767 tokenpak serve
export GOOGLE_GEMINI_BASE_URL="http://localhost:8767"
Auth errors — 401 or invalid API key¶
Gemini CLI still needs a valid Gemini API key. TokenPak does not replace credentials. Confirm:
GEMINI_API_KEYis exported in the same shell that runsgemini.- The key is valid for Google AI Studio / Gemini API.
- You are not mixing Vertex AI variables with Google AI Studio variables.
Environment caching — variable changes do not take effect¶
Gemini CLI reads environment variables when the process starts. After changing GOOGLE_GEMINI_BASE_URL:
- Stop the current
geminiprocess. - Export the new value.
- Start a new
geminicommand. - Check
tokenpak statusagain.
If you run Gemini CLI from an editor task runner or terminal multiplexer, make sure that runner inherits the updated environment.
Tools or function-calling requests fail¶
Tool and function-calling requests are not verified by current release evidence for this integration. For workflows that need verified behavior, use a first-class integration (Claude Code or Codex) or a tested adapter (OpenAI SDK, Anthropic SDK or LiteLLM).