Enterprise inference

Same workload.A smaller token bill.

Pick a model. Set your throughput. We'll profile your workload for a lower-cost serving configuration.

Workload profileServing plan

01 / Traffic

Representative requests

8K IN650 OUT

02 / Cached context

Route repeated prefixes together

03 / Matched compute

Prefill

PROMPT PHASE

Decode

OUTPUT PHASE

04 / Output SLA

Sustained generation target

TPS

DEFAULT

Example profiling path. The selected approach depends on model and workload behavior.

Model your inference costs.

Compare published token rates against your workload and cache profile.

Choose a model USD per 1M tokens

Across all requests, not one response.

01K1M1B

Eligible input only. Example, not a forecast.

50%
0%100%
Assumes 4:1 input/output · 24h/day · 30 days · edit assumptions
Optional service target

Public-rate estimate 50% cache scenario

$16,070

per month · $192,845 annualized

Reference · 0% input cache hits$18,144
Scenario · 50% input cache hits$16,070
Fresh inputCached inputOutput

$2,074 lower per month in this scenario

Keep the model. Lower the cost.

Final pricing follows your workload benchmark.

DeepInfra via OpenRouter · 2026-09-06 · Rates & assumptions

USD per 1M tokens: $0.09 fresh input · $0.05 cached input · $0.34 output. Route: deepinfra/turbo, FP4. Model pricing · Provider rate source

Costs use aggregate output TPS × active hours × 3,600 × 30 days, plus input at your selected ratio. Annualized values use twelve 30-day months. Both scenarios use the same provider’s rates and output volume. Bar segments represent fresh input (gray), cached input (hatched), and output (black).

OpenRouter publishes a cache-read rate for this provider endpoint. Its endpoint metadata reported implicit caching as unavailable at retrieval. Do not assume a cache hit. Apply this rate only to input confirmed eligible by the provider's cache behavior and request configuration. Cache hits are an editable assumption, not a prediction. No separate cache-write rate is published. Additional cache-write, storage, network, platform fees, and taxes are excluded, not assumed free.

Market-reference provider route, not an OpenRelay availability or price quote. These token-charge scenarios are not your actual bill, measured savings, or an SLA. Final pricing and service targets require a workload benchmark. sales@openrelay.inc

Choose the outcome

Start with the SLA that matters.

Not sure which to pick? Start with tokens per second. Let us profile your workload and lower the bill.

Recommended

Default SLA

TPS

Tokens per second

Set a sustained output rate for the workload you actually run.

Best starting point

Interactive

TTFT + ITL

Latency

Set response targets for chat, copilots, and user-facing agents.

User experience

Scheduled

TOKENS / WINDOW

Batch deadline

Set a token volume and the time by which it must finish.

Completion time

What we profile

Where we look for savings.

Approaches benchmarked per workload

01

Context reuse

Test cache-aware routing and shared-prefix reuse.

02

Kernel + batch tuning

Benchmark kernels, quantization, and batch shape.

03

Fit the hardware

Match prefill and decode phases to measured demand.

Bring the workload

Let's find the expensive tokens.

Send your model, traffic shape, and target SLA. We will define a representative benchmark.

Email sales@openrelay.inc

Final pricing and SLA follow a workload benchmark.