Models ready to run

Choose a model, call the API, and pay only for the tokens you use.

Most popularOpenAI

GPT-OSS 120B

$0.15 / $0.60

input / output per 1M

OpenAI's 120B open-weight model, with reasoning, tool calling, and structured output support.

Text · Reasoning · Code

Model page
Best valueOpenAI

GPT-OSS 20B

$0.05 / $0.20

input / output per 1M

The smaller GPT-OSS variant, with tool calling and structured output support at lower token rates.

Text · Code

Model page
Best valueGoogle

Gemma 4 31B NVFP4 32K

$0.081 / $0.306

input / output per 1M

An NVFP4-quantized Gemma 4 31B deployment with a 32K-token context window and lower token rates.

Text · Vision · Reasoning

Model page
Google

Gemma 4 31B

$0.14 / $0.40

input / output per 1M

Google's dense 31B model with text and image input, reasoning, tool calling, and a 32K-token context window.

Text · Vision · Reasoning

Model page
Zhipu

GLM 5.2

$1.82 / $5.72

input / output per 1M

Zhipu's GLM model for reasoning, coding, and tool use in English and Chinese, with a 1M-token context window.

Text · Reasoning · Code

Model page
Zhipu

GLM 5.3 Flash

$0.195 / $0.65

input / output per 1M

Zhipu's natively multimodal GLM 5.3 Flash, served on our own H100s. Hybrid linear and sparse attention keeps long-context serving cheap, and it takes text or images in.

Text · Vision · Reasoning · Code

DeepSeek

DeepSeek-OCR 2

$0.039 / $0.039

input / output per 1M

DeepSeek's second-generation OCR model. Reads document images (scans, receipts, screenshots, tables) and returns structured markdown that preserves headings, tables, and layout.

Vision

Model page
DeepSeek

DeepSeek V3.1 Terminus

$0.27 / $1.00

input / output per 1M

DeepSeek's V3.1 Terminus release: a sparse mixture-of-experts model with a hybrid thinking mode, tuned for multi-step agentic work and consistent tool calling.

Text · Reasoning · Code

Model page

Don't see the model you need?

Need a guaranteed SLA for a model?

The pay-per-token catalog uses shared capacity. Contact us for dedicated capacity with contractual availability and performance targets.

Availability
Uptime terms and service credits are agreed in the contract for dedicated capacity.
Performance
Time-to-first-token and throughput targets are sized to the selected model and traffic profile.
Support
Support channels and response times are defined as part of the deployment agreement.

Need private inference?

Confidential inference: your request runs in a hardware-isolated, cryptographically attested environment

Run supported models in hardware-isolated environments with cryptographic attestation, so you can verify the environment handling your requests.

  • Hardware isolated
  • Cryptographically attested
  • Verifiable execution

Confidential computing powered by Caution

Start calling models in minutes

Grab an API key, pick a model, and send your first request.