System One (typed decisions)

Last updated:

On this page
text
POST https://api.staik.se/v1/systemone

Most decisions in an agent or workflow do not need generated text: is this email phishing, which queue does this ticket belong in, is the evidence sufficient to answer? You send a state (text or JSON) and one or more typed questions, and get back a probability for every option, never free text. The answer always matches your schema, so there is nothing to parse or validate.

The format is compatible with TypeSafe's Jev API. An existing Jev client works by setting base_url to https://api.staik.se and using your staik key. Your data never leaves Sweden.

bash
curl https://api.staik.se/v1/systemone \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-st-your-key" \
  -d '{
    "state": {
      "subject": "Your account has been suspended",
      "body": "Log in via the link and confirm your banking credentials within 24 hours."
    },
    "questions": {
      "category": {
        "type": "choice",
        "instructions": "What kind of email is this?",
        "criteria": {
          "support": "Customer service, ticket, complaint",
          "billing": "Invoice, payment, receipt",
          "threat": "Phishing, fraud, social engineering"
        }
      },
      "phishing": {
        "type": "noul",
        "instructions": "Is the sender trying to obtain credentials or money?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this for the recipient?",
        "criteria": ["Not at all", "Low", "High", "Critical"]
      }
    }
  }'

Question types

TypecriteriaAnswer
choiceObject {name: description}, 2–20 optionschoice (chosen name), probabilities per name, confidence
noulOptional {"true": ..., "false": ...}noul = probability of yes (0–1)
scoreList of ordered levels, 2–20score = expected level index, probabilities, legend, confidence

instructions is the question in plain language. Both state and instructions can be text or JSON; field names are kept as labels in what the model reads.

Response

json
{
  "request_id": "7f3c…",
  "model": "gemma4:31b",
  "answers": {
    "category": {
      "type": "choice",
      "choice": "threat",
      "confidence": 0.94,
      "probabilities": { "support": 0.02, "billing": 0.02, "threat": 0.96 }
    },
    "phishing": { "type": "noul", "noul": 0.97 },
    "urgency": {
      "type": "score",
      "score": 2.1,
      "legend": { "0": "Not at all", "1": "Low", "2": "High", "3": "Critical" },
      "probabilities": { "0": 0.03, "1": 0.12, "2": 0.57, "3": 0.28 },
      "confidence": 0.61
    }
  },
  "usage": { "input_tokens": 212, "output_tokens": 0 },
  "latency_ms": 184
}

confidence is 0 for a uniform distribution and 1 when all probability sits on a single option. request_id is also returned in the x-typesafe-request-id header.

Models

modelStrength
gemma4:31b (default)Highest accuracy, and fastest when several questions share one state
qwen3.6:35b-a3bBest at signalling uncertainty on factual questions; pick it when automating "accept if confident, otherwise escalate" on knowledge questions
qwen3.5:9bSmallest, slightly lower accuracy

Without model, or with a name we do not recognise (such as a Jev client's jev-latest), gemma4:31b is used. Other models return 422.

Calibration

Probabilities are read directly from the model's next-token distribution and scaled by a per-model temperature fitted on public evaluation data. On a separate test set (Kev's transfer-v4, 764 questions) the calibration error (ECE) is 0.03–0.04 for all three models: when the model says 80 %, it is right roughly 80 % of the time. Calibration is shared across all customers. Measure against your own examples before automating decisions that are expensive to get wrong.

Cost and limits

  • Billed against the same token quota as chat. The state counts once per request, no matter how many questions you ask. Output is not billed.
  • At most 20 options per question and 32 questions per request.
  • A request takes one slot in the same queue as your chat requests.

Ask several questions about the same state

Questions run in parallel and share the processed state, so three questions in one request are both faster and cheaper than three separate requests.

Errors

StatusCause
400Invalid JSON, or the state does not fit the model's context window
401Missing or invalid API key
422Invalid question (unknown type, too many options, missing state) or unsupported model
429Quota exhausted or queue full (see Retry-After)
502 / 503Temporary model backend error, retry with backoff