Same models.
Lower API price.
Keep the AI stack you already use. Relay changes the economics underneath it, through one OpenAI-compatible API.
Keep your stack.
Change the economics underneath.
One OpenAI-compatible API. Same models, same prompts and the same workflows. Route only the workloads where Relay gives you better economics.
- Same models
- Lower API pricing
- No forced migration
Don’t take our word for it.
See what the economics look like across real AI agent workloads.
Cost per successful AI agent task, not per API request. Total spend includes all model usage, including unsuccessful runs. DeepSeek and Sonnet achieved virtually identical measured pass rates on the same workload: 98.4% and 98.0%.
GLM had a lower published token price than Sonnet, yet cost roughly twice as much per successful AI agent task: $1.59 versus $0.79.
The difference was verbosity, not quality. GLM produced an average of 375 output tokens per task versus Sonnet’s 56, nearly 7× as many, while their measured pass rates were essentially identical.
96% of GLM’s output tokens were reasoning tokens. Sonnet used none. Those extra tokens still carry a cost.
Price per million tokens tells you what a token costs. It doesn’t tell you what the agent task costs.Change base_url.
Nothing else.
Same models. Same prompts. Same workflows.
- Keep your existing stack
- Keep your model choice
- Move traffic selectively
client = OpenAI(
api_key = RELAY_API_KEY,
base_url = "https://relaygpu.com/v2/openai/v1",
)At production scale, this compounds.
Small unit-cost differences become six-figure annual savings.
250M successful tasks a year, measured against Sonnet.
5B generated characters a year, against list pricing.
5M image renders a year. Video routes save separately, below.
Every comparison uses a defined unit and an explicit comparator, with same-model pricing where applicable. Annual figures are volume-scaled examples, not invoices. Pricing advantages vary by model and provider availability, and Relay only claims savings where current pricing supports them.



| Media route | Comparator | Relay | Scale | Saving |
|---|---|---|---|---|
| Gemini image | $0.077 / image | $0.046 / image | 5M images | $155K |
| HappyHorse 1.1 | $0.168 / sec | $0.1512 / sec | 1M x 5 sec | $84K |
| SeeDance 2.5 | $21.40 / 1M video tokens | $10.01 / 1M video tokens | Route-level | 53% less |
| SeeDance Mini | $7.00 / 1M tokens | $3.22 / 1M tokens | Route-level | 54% less |
How Relay lowers the price.
The advantage sits in the supply economics, not in changing the model you use.
Direct provider relationships
Relay works directly with infrastructure and technology providers to negotiate commercial rates.
Commercial sourcing
Multiple provider relationships create more pricing options than any single customer reaches alone.
Infrastructure optionality
Equivalent model capacity can be sourced through different infrastructure routes where available.
Prove it before you move it.
We’re confident enough in Relay’s economics to fund the benchmark.
We’ll give your team free Relay credits to test your existing workloads against Relay before you commit. Compare cost, TTFT, latency, reliability and model performance against what you use today, measured from where your traffic actually runs.
If Relay wins, move the workloads that benefit. If it doesn’t, don’t.
Get free Relay creditsBook a call and we’ll run it with youPrice your stack before launch.
Tell us the models you are considering and your expected token, voice or media volume. Compare the economics before committing to anything.
Compare what you pay today.
Keep your existing models, application and workflows. Move only the workloads where Relay produces better economics.