Laya, decisions instead of generated text
Laya is an encoder-only model from Convai Innovations: send a state and typed questions, get back labels, scores and probabilities. It generates nothing, so only input tokens bill.
- Choice, score, yes or no
- English and multilingual
- Input-only billing
- Apache 2.0
Per token, nothing else.
Input / 1M tokens
$0.02
Decisions meter input tokens only. Nothing is generated, so there is no output rate.
openrelay/layaFull catalog on the inference pricing page. Add a card and make a deposit to start.
import os
import requests
API_KEY = os.environ["OPENRELAY_API_KEY"] # an or_ key
resp = requests.post(
"https://inference.openrelay.inc/v1/decisions",
headers={"Authorization": f"Bearer {API_KEY}"},
json={
"model": "openrelay/laya",
"state": "We were billed twice for March. Refund it or we cancel.",
"questions": {
"department": {"type": "choice", "instructions": "Which team handles this?",
"criteria": ["billing", "technical", "other"]},
"churn_risk": {"type": "noul",
"instructions": "Does the user threaten to leave?"},
},
},
)
print(resp.json()["answers"])The spec sheet.
What the model is, what it takes in, and the surface it serves on.
- Model id
- openrelay/laya
- Type
- Encoder-only decision model: it labels and scores, it does not generate text
- Question types
- choice (a label from your options), score (a level on your scale), noul (the probability of yes)
- Checkpoints
- English (421M parameters) and multilingual, picked per request from the script and language of the state
- Limits
- 64 questions per request, 8,192 tokens read per question, 32 KB of state
- Endpoint
- POST /v1/decisions, online only (the Batch API does not accept it)
- License
- Apache 2.0, open weights
Where Laya earns its place.
What this model is actually for, versus the rest of the catalog.
Typed answers, nothing to parse
Each answer is a field: a label with its probabilities, a score, or the probability of yes. There is no generated text to parse or validate.
One read per question
The state is read once per question and only input tokens bill. Ten questions about one ticket cost ten reads, in one round trip.
Measure before you route
The served checkpoints are Convai's base models, which its authors report near chance zero-shot on their typed-decisions benchmark, with uncalibrated probabilities. Validate on a labeled sample and set thresholds from it.
Laya, answered.
What is Laya?
An encoder-only decision model by Convai Innovations, released under Apache 2.0. It reads a state and answers typed questions with a label or a score plus probabilities, in one forward pass.
How is Laya billed?
Per input token, with no output rate. usage.prompt_tokens counts the tokens read, summed over questions, and the state is read once per question. Tokens past the per-question limit are not read and not billed.
Which languages does it handle?
Each request is routed by the script and language of its state: English to the English checkpoint, other languages and non-Latin scripts to the multilingual one. routing.checkpoint in the response names the one that answered.
Can I run decisions through the Batch API?
No. /v1/decisions is online only, and the Batch API does not accept it.
First request in five lines.
POST to inference.openrelay.inc/v1/decisions and run Laya per input token. No contract, no minimums.