System One (typed decisions)
Last updated:
POST https://api.staik.se/v1/systemoneMost decisions in an agent or workflow do not need generated text: is this email phishing, which queue does this ticket belong in, is the evidence sufficient to answer? You send a state (text or JSON) and one or more typed questions, and get back a probability for every option, never free text. The answer always matches your schema, so there is nothing to parse or validate.
The format is compatible with TypeSafe's Jev API. An existing Jev client works by
setting base_url to https://api.staik.se and using your staik key. Your data
never leaves Sweden.
curl https://api.staik.se/v1/systemone \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-st-your-key" \
-d '{
"state": {
"subject": "Your account has been suspended",
"body": "Log in via the link and confirm your banking credentials within 24 hours."
},
"questions": {
"category": {
"type": "choice",
"instructions": "What kind of email is this?",
"criteria": {
"support": "Customer service, ticket, complaint",
"billing": "Invoice, payment, receipt",
"threat": "Phishing, fraud, social engineering"
}
},
"phishing": {
"type": "noul",
"instructions": "Is the sender trying to obtain credentials or money?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this for the recipient?",
"criteria": ["Not at all", "Low", "High", "Critical"]
}
}
}'Question types
| Type | criteria | Answer |
|---|---|---|
choice | Object {name: description}, 2–20 options | choice (chosen name), probabilities per name, confidence |
noul | Optional {"true": ..., "false": ...} | noul = probability of yes (0–1) |
score | List of ordered levels, 2–20 | score = expected level index, probabilities, legend, confidence |
instructions is the question in plain language. Both state and instructions
can be text or JSON; field names are kept as labels in what the model reads.
Response
{
"request_id": "7f3c…",
"model": "gemma4:31b",
"answers": {
"category": {
"type": "choice",
"choice": "threat",
"confidence": 0.94,
"probabilities": { "support": 0.02, "billing": 0.02, "threat": 0.96 }
},
"phishing": { "type": "noul", "noul": 0.97 },
"urgency": {
"type": "score",
"score": 2.1,
"legend": { "0": "Not at all", "1": "Low", "2": "High", "3": "Critical" },
"probabilities": { "0": 0.03, "1": 0.12, "2": 0.57, "3": 0.28 },
"confidence": 0.61
}
},
"usage": { "input_tokens": 212, "output_tokens": 0 },
"latency_ms": 184
}confidence is 0 for a uniform distribution and 1 when all probability sits on a
single option. request_id is also returned in the x-typesafe-request-id header.
Models
model | Strength |
|---|---|
gemma4:31b (default) | Highest accuracy, and fastest when several questions share one state |
qwen3.6:35b-a3b | Best at signalling uncertainty on factual questions; pick it when automating "accept if confident, otherwise escalate" on knowledge questions |
qwen3.5:9b | Smallest, slightly lower accuracy |
Without model, or with a name we do not recognise (such as a Jev client's
jev-latest), gemma4:31b is used. Other models return 422.
Calibration
Probabilities are read directly from the model's next-token distribution and scaled by a per-model temperature fitted on public evaluation data. On a separate test set (Kev's transfer-v4, 764 questions) the calibration error (ECE) is 0.03–0.04 for all three models: when the model says 80 %, it is right roughly 80 % of the time. Calibration is shared across all customers. Measure against your own examples before automating decisions that are expensive to get wrong.
Cost and limits
- Billed against the same token quota as chat. The state counts once per request, no matter how many questions you ask. Output is not billed.
- At most 20 options per question and 32 questions per request.
- A request takes one slot in the same queue as your chat requests.
Ask several questions about the same state
Questions run in parallel and share the processed state, so three questions in one request are both faster and cheaper than three separate requests.
Errors
| Status | Cause |
|---|---|
400 | Invalid JSON, or the state does not fit the model's context window |
401 | Missing or invalid API key |
422 | Invalid question (unknown type, too many options, missing state) or unsupported model |
429 | Quota exhausted or queue full (see Retry-After) |
502 / 503 | Temporary model backend error, retry with backoff |