Models/ Google

Gemma 4 31B, the open model that reads images

Google's Gemma 4 31B takes images and text in one request: screenshots, scanned pages, charts, photos. Rates are on the pricing table below, straight from the catalog.

  • Vision + text input
  • 256K context
  • Batch at 50% off

Per token, nothing else.

Input / output per 1M tokens

$0.99 / $1.49

Batch jobs: $0.495 / $0.745 (50% off, 24h window)

openrelay/gemma-4-31b

Full catalog on the inference pricing page. Add a card and make a deposit to start.

gemma.py · vision inputpython
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.openrelay.inc/v1",
    api_key="or_••••••••",          # same SDK, new base URL
)

import base64
img = base64.b64encode(open("chart.png", "rb").read()).decode()

resp = client.chat.completions.create(
    model="openrelay/gemma-4-31b",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What does this chart imply about Q3?"},
        {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img}"}},
    ]}],
)
print(resp.choices[0].message.content)

The spec sheet.

What the model is, what it takes in, and the surface it serves on.

Model id
openrelay/gemma-4-31b
Parameters
31B (open weights)
Modalities
Text + image in, text out
Context
256K tokens online, 128K in batch
Endpoint
POST /v1/chat/completions with optional image parts

Where Gemma 4 31B earns its place.

What this model is actually for, versus the rest of the catalog.

Vision where text models stop

Classify scanned pages by layout, moderate image posts, read charts, describe screenshots. Anything where the signal is in pixels routes here while text-only work stays on cheaper models.

Batch at half the online rate

Moderation, classification, and summarization backlogs run through the Batch API at 50% of the per-token rate on the standard tier, up to 128K tokens a request. A batch runs one model, so image and text records share a Gemma batch; a cheaper text model runs through the online API.

Whole documents in one request

256K tokens online and 128K in batch fit most transcripts, contracts, and reports whole, so summarization and analysis run one request per document. That is how the batch summarization recipe uses it.

Run it in batch at half price.

Gemma runs the per-document step of the batch summarization recipe and the image records inside moderation and classification jobs, at half the online rate on the standard tier.

Batch Inference API overview

Gemma 4 31B, answered.

Is there a hosted Gemma API?

Yes. Gemma 4 31B is live behind the OpenAI-compatible endpoint at inference.openrelay.inc/v1, vision input included, with the standard OpenAI SDK. No Google Cloud project required.

What does the Gemma API cost?

The pricing table above carries the current per-token rate, read from the catalog. Standard batch halves it; expedited batch bills the online rate. If most of your records are text only, send those to GPT-OSS 20B, which costs less, and route only the records with images to Gemma.

How long a document can Gemma read in one request?

256K tokens through the online API and 128K in batch, which covers most transcripts, contracts, and reports whole. For a longer input, or for merging many summaries into one, GLM 5.2 takes 1M tokens.

Can Gemma read scanned documents?

Yes, send page images as message parts. For pure transcription at archive scale, DeepSeek-OCR 2 is cheaper per page; Gemma earns its rate when pages need interpretation: forms with implicit structure, charts, handwriting mixed with print.

First request in five lines.

Point the OpenAI SDK at inference.openrelay.inc/v1 and run Gemma 4 31B per token. No contract, no minimums.