Relay for production AI

Same models.
Lower API price.

Keep the AI stack you already use. Relay changes the economics underneath it, through one OpenAI-compatible API.

Agents91%lower cost per successful AI agent task, measured against a frontier model
Voice51%lower unit cost per 1,000 generated characters
Media53%lower route cost on the same model
Low-friction adoption

Keep your stack.
Change the economics underneath.

One OpenAI-compatible API. Same models, same prompts and the same workflows. Route only the workloads where Relay gives you better economics.

  • Same models
  • Lower API pricing
  • No forced migration

Don’t take our word for it.

See what the economics look like across real AI agent workloads.

$0.07per 1,000 successful tasks
against
$0.79same AI agent workload, frontier model
91%lower in the measured comparison
$0.07
DeepSeek
$0.49
GPT
$0.79
Sonnet
$1.29
Qwen
$1.59
GLM
Cost per 1,000 successful AI agent tasks · lower is better

Cost per successful AI agent task, not per API request. Total spend includes all model usage, including unsuccessful runs. DeepSeek and Sonnet achieved virtually identical measured pass rates on the same workload: 98.4% and 98.0%.

Token price isn’t task price

GLM had a lower published token price than Sonnet, yet cost roughly twice as much per successful AI agent task: $1.59 versus $0.79.

The difference was verbosity, not quality. GLM produced an average of 375 output tokens per task versus Sonnet’s 56, nearly 7× as many, while their measured pass rates were essentially identical.

96% of GLM’s output tokens were reasoning tokens. Sonnet used none. Those extra tokens still carry a cost.

Price per million tokens tells you what a token costs. It doesn’t tell you what the agent task costs.
<1sDirect mode latency
5models, same workload and scoring

Change base_url.
Nothing else.

Same models. Same prompts. Same workflows.

  • Keep your existing stack
  • Keep your model choice
  • Move traffic selectively
python
client = OpenAI(
    api_key  = RELAY_API_KEY,
    base_url = "https://relaygpu.com/v2/openai/v1",
)

At production scale, this compounds.

Small unit-cost differences become six-figure annual savings.

$180KAgents

250M successful tasks a year, measured against Sonnet.

$255KVoice

5B generated characters a year, against list pricing.

$155KMedia

5M image renders a year. Video routes save separately, below.

Every comparison uses a defined unit and an explicit comparator, with same-model pricing where applicable. Annual figures are volume-scaled examples, not invoices. Pricing advantages vary by model and provider availability, and Relay only claims savings where current pricing supports them.

Image, at 5M renders$155K$0.077 to $0.046 per image
Video, at 1M renders$84K5 sec HappyHorse route
SeeDance 2.553%$21.40 to $10.01 per 1M video tokens
Media routeComparatorRelayScaleSaving
Gemini image$0.077 / image$0.046 / image5M images$155K
HappyHorse 1.1$0.168 / sec$0.1512 / sec1M x 5 sec$84K
SeeDance 2.5$21.40 / 1M video tokens$10.01 / 1M video tokensRoute-level53% less
SeeDance Mini$7.00 / 1M tokens$3.22 / 1M tokensRoute-level54% less

How Relay lowers the price.

The advantage sits in the supply economics, not in changing the model you use.

01

Direct provider relationships

Relay works directly with infrastructure and technology providers to negotiate commercial rates.

02

Commercial sourcing

Multiple provider relationships create more pricing options than any single customer reaches alone.

03

Infrastructure optionality

Equivalent model capacity can be sourced through different infrastructure routes where available.

Providersnegotiated rates
Relaysame model, lower price
Your applicationunchanged
Put Relay to the test

Prove it before you move it.

We’re confident enough in Relay’s economics to fund the benchmark.

We’ll give your team free Relay credits to test your existing workloads against Relay before you commit. Compare cost, TTFT, latency, reliability and model performance against what you use today, measured from where your traffic actually runs.

If Relay wins, move the workloads that benefit. If it doesn’t, don’t.

Get free Relay creditsBook a call and we’ll run it with you
Building?

Price your stack before launch.

Tell us the models you are considering and your expected token, voice or media volume. Compare the economics before committing to anything.

Already live?

Compare what you pay today.

Keep your existing models, application and workflows. Move only the workloads where Relay produces better economics.

Relay economics
1 Send usage2 We compare3 You decide

No forced migration.

Visit relaygpu.com