Default SLA
TPS
Tokens per second
Set a sustained output rate for the workload you actually run.
Enterprise inference
Pick a model. Set your throughput. We'll profile your workload for a lower-cost serving configuration.
01 / Traffic
Representative requests
02 / Cached context
Route repeated prefixes together
03 / Matched compute
Prefill
PROMPT PHASE
Decode
OUTPUT PHASE
04 / Output SLA
Sustained generation target
DEFAULT
Example profiling path. The selected approach depends on model and workload behavior.
Compare published token rates against your workload and cache profile.
Choose a model USD per 1M tokens
Across all requests, not one response.
Eligible input only. Example, not a forecast.
Public-rate estimate 50% cache scenario
$16,070
per month · $192,845 annualized
$2,074 lower per month in this scenario
Keep the model. Lower the cost.
Final pricing follows your workload benchmark.
USD per 1M tokens: $0.09 fresh input · $0.05 cached input · $0.34 output. Route: deepinfra/turbo, FP4. Model pricing · Provider rate source
Costs use aggregate output TPS × active hours × 3,600 × 30 days, plus input at your selected ratio. Annualized values use twelve 30-day months. Both scenarios use the same provider’s rates and output volume. Bar segments represent fresh input (gray), cached input (hatched), and output (black).
OpenRouter publishes a cache-read rate for this provider endpoint. Its endpoint metadata reported implicit caching as unavailable at retrieval. Do not assume a cache hit. Apply this rate only to input confirmed eligible by the provider's cache behavior and request configuration. Cache hits are an editable assumption, not a prediction. No separate cache-write rate is published. Additional cache-write, storage, network, platform fees, and taxes are excluded, not assumed free.
Market-reference provider route, not an OpenRelay availability or price quote. These token-charge scenarios are not your actual bill, measured savings, or an SLA. Final pricing and service targets require a workload benchmark. sales@openrelay.inc
Choose the outcome
Not sure which to pick? Start with tokens per second. Let us profile your workload and lower the bill.
Default SLA
TPS
Set a sustained output rate for the workload you actually run.
Interactive
TTFT + ITL
Set response targets for chat, copilots, and user-facing agents.
Scheduled
TOKENS / WINDOW
Set a token volume and the time by which it must finish.
What we profile
Approaches benchmarked per workload
Test cache-aware routing and shared-prefix reuse.
Benchmark kernels, quantization, and batch shape.
Match prefill and decode phases to measured demand.
Bring the workload
Send your model, traffic shape, and target SLA. We will define a representative benchmark.
Email sales@openrelay.incFinal pricing and SLA follow a workload benchmark.