GPT-6 Astra
−85%
Below the model's base public API rates.
- Input
- $1.50$10.00
- Output
- $7.50$50.00
USD per 1 million tokens
No long-context surcharge.
Undercut the price, not the model.
An OpenAI- and Anthropic-compatible gateway for Claude Code and production apps. Choose the right model, pay only for tokens used, and track every request.
$ export ANTHROPIC_BASE_URL=https://undercut.pro
$ export ANTHROPIC_API_KEY=gw_••••••••••••
$ claude
✓ Connected through UndercutAIModel routing
Switch models with one parameter. Context windows up to 1 million tokens; Grok supports up to 600,000 tokens. No new SDK or billing account.
Deep reasoning
Balanced & fast
Advanced reasoning
Frontier intelligence
Standard UndercutAI rates
Our standard discounts are 55% or more off the listed base public API rates: 85% for GPT-6 Astra and 80% for Claude Fable 5.1. No subscription or limited-time offer. Compare matching token categories; the prices below are our applied rates.
Permanent prices for short and long context.
−85%
Below the model's base public API rates.
USD per 1 million tokens
No long-context surcharge.
−80%
Below the model's base public API rates.
USD per 1 million tokens
No long-context surcharge.
| Model | Standard tokens | Cached tokens | Savings | ||
|---|---|---|---|---|---|
| Input | Output | Cache read | Cache write | ||
Claude Opus 5 claude-opus-5 | $5.00$2.25 | $25.00$11.25 | $0.50$0.225−55% vs the provider's cache-read rate | $6.25$2.8125Reference: 5-minute cache write | −55% |
Claude Sonnet 5 claude-sonnet-5 | $2.00$0.90 | $10.00$4.50 | $0.20$0.09−55% vs the provider's cache-read rate | $2.50$1.125Reference: 5-minute cache write | −55% |
Claude Fable 5 claude-fable-5 | $10.00$4.50 | $50.00$22.50 | $1.00$0.45−55% vs the provider's cache-read rate | $12.50$5.625Reference: 5-minute cache write | −55% |
Claude Fable 5.1Special rate claude-fable-5.1 | $10.00$2.00 | $50.00$10.00 | $0.25$0.05−80% vs the provider's cache-read rate | $12.50$2.50Reference: 5-minute cache write | −80% |
GPT-6 AstraSpecial rate gpt-6-astra | $10.00$1.50 | $50.00$7.50 | $1.00$0.15−85% vs the provider's cache-read rate | $12.50$1.875 | −85% |
GPT-5.6 Sol gpt-5.6-sol | $4.00$1.80 | $20.00$9.00 | $0.40$0.18−55% vs the provider's cache-read rate | $5.00$2.25 | −55% |
GPT-5.6 Terra gpt-5.6-terra | $2.00$0.90 | $12.00$5.40 | $0.20$0.09−55% vs the provider's cache-read rate | $2.50$1.125 | −55% |
GPT-5.6 Luna gpt-5.6-luna | $0.20$0.09 | $1.20$0.54 | $0.02$0.009−55% vs the provider's cache-read rate | $0.25$0.1125 | −55% |
Gemini 3.1 Pro gemini-3.1-pro-preview | $2.00$0.90 | $12.00$5.40 | $0.20$0.09−55% vs the provider's cache-read rate | No separate public write rate$0.90UndercutAI: at the input rate | −55% |
Gemini 3.7 Flash gemini-3.7-flash | $0.75$0.3375 | $3.75$1.6875 | $0.075$0.03375−55% vs the provider's cache-read rate | No separate public write rate$0.3375UndercutAI: at the input rate | −55% |
Gemini 3.8 Flash gemini-3.8-flash | $0.75$0.3375 | $3.75$1.6875 | $0.075$0.03375−55% vs the provider's cache-read rate | No separate public write rate$0.3375UndercutAI: at the input rate | −55% |
Grok 4.6 grok-4.6 | $2.00$0.90 | $6.00$2.70 | $0.50$0.225−55% vs the provider's cache-read rate | No separate public write rate$0.90UndercutAI: at the input rate | −55% |
Kimi K3 kimi-k3 | $3.00$1.35 | $15.00$6.75 | $0.30$0.135−55% vs the provider's cache-read rate | No separate public write rate$1.35UndercutAI: at the input rate | −55% |
Reference prices reviewed: .
The discount compares our rates with providers' base public text-token rates, excluding their long-context surcharges. This defines the price comparison, not our context limits. Claude cache writes use the 5-minute reference. UndercutAI applies the listed flat token rates without a long-context surcharge or a separate time-based cache-storage charge; managed cache retention is not offered by this table.
The reference includes OpenAI Sol promotional pricing available at least through November 21, 2026, and Gemini 3.7/3.8 Flash promotional pricing through December 31, 2026. These provider promotions are separate from UndercutAI's permanent discounts of at least 55%.
Built for developers
Keep your OpenAI or Anthropic client. Change the base URL and API key, then monitor model, input tokens, output tokens, and derived cost from one dashboard.
USD balance with request-level spend history.
Create and revoke credentials instantly.
Configured per-million input and output rates.
Use a custom Anthropic-compatible endpoint.
Current balance
$48.27
claude-sonnet-5
12,480 tokens
$0.084
gpt-5.6-sol
7,234 tokens
$0.037
claude-opus-5
3,902 tokens
$0.126
Frequently asked questions
Clear answers about pricing, model authenticity, verification, and what stays private to protect the service.
A practical trust policy
We do not treat a model's self-reported name as proof. Controlled, repeatable comparisons are more meaningful than screenshots or unverifiable claims.
UndercutAI is a prepaid AI API gateway. One key and one balance give you access to supported frontier models through familiar OpenAI- and Anthropic-compatible endpoints. You choose the model in every request and see token-level usage and cost in one dashboard.
UndercutAI sets its own retail rates. Our standard discounts start at 55% below the listed public reference rates and vary by model; they do not describe our upstream costs. Cache hits reduce the cost of repeated input separately, not the cost of output. A lower price does not permit us to substitute the model you request.
No. Your USD balance pays for usage within UndercutAI at our published rates; it is not transferred to an Anthropic or OpenAI account. Each model's discount is already included in token prices, not added as a credit multiplier on top-up. The dashboard shows the amount credited and the available payment methods; currency conversion and payment fees may affect what you pay.
No. Discounts of at least 55% are our permanent standard pricing policy relative to the listed base public API rates, with no scheduled expiry. GPT-6 Astra has an 85% discount and Claude Fable 5.1 has an 80% discount. Dollar token prices are not fixed forever: they may change when providers change their reference rates. The pricing table shows the rates currently applied.
No. UndercutAI does not intentionally substitute a requested model with an unrelated lower-cost model. That is the whole point of "Undercut the price, not the model." — the model identifier in your request is the product you are buying. Routing details may remain confidential, but a lower price is not permission to silently downgrade the model.
A model's self-description and the model field in an API response are not independent proof of its identity. Controlled side-by-side comparisons can help assess consistency: use the same prompts and supported settings through UndercutAI and the official API, then compare capabilities, tool calls, long-context recall and token usage over several runs. These comparisons are useful evidence, not cryptographic proof.
No. Self-identification is weak evidence: a system prompt can change the answer, and different models can imitate the same response. Capability tests with fixed inputs, low temperature, repeat runs, and an official baseline are much harder to fake and produce a more useful comparison.
Publishing credentials, account identifiers, raw upstream headers, or routing topology would make the infrastructure easy to abuse or copy and could reduce availability for every customer. We keep those operational details private while exposing the information customers need to audit their own usage: requested model, token categories, applied rate, cost, and history.
Model output is probabilistic. Temperature, system instructions, tool definitions, context, provider-side model updates, and even repeated runs can change an answer. For a fair comparison, keep every input and parameter identical, run several samples, and judge capability and consistency rather than exact wording.
Usually only the base URL and API key. Keep your existing OpenAI or Anthropic client, select a supported model, and monitor usage from the dashboard. The quickstart page includes ready-to-copy examples for popular SDKs and developer tools.
Standard rates · save at least 55%
Create an account, add funds, and issue your first key in minutes.
Create free account