Responses (OpenAI)
Last updated:
POST https://api.staik.se/v1/responsesEndpoint compatible with OpenAI's Responses API. The OpenAI SDKs'
responses.create, Codex CLI and other clients that speak the Responses protocol
work directly. Authenticate with Authorization: Bearer
(same key).
Requests run on the same path as chat completions:
same models, quotas, queues, failover and billing. Usage is recorded under the
endpoint responses.
from openai import OpenAI
client = OpenAI(api_key="sk-st-your-key", base_url="https://api.staik.se/v1")
response = client.responses.create(
model="qwen3.6:35b-a3b",
instructions="Answer briefly.",
input="Explain how a transformer works.",
)
print(response.output_text)Stateless: send the full history
staik never stores responses. Every request must contain the full conversation in
input, including earlier model output, tool calls and tool results. store is
ignored and the response always has "store": false.
Function tools
Tools use the Responses format (type: "function" with name and parameters
directly on the tool). The model answers with function_call items, you run the
tool and send the result back as function_call_output with the same call_id:
import json
tools = [{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
}]
history = [{"role": "user", "content": "What's the weather in Gothenburg?"}]
response = client.responses.create(model="qwen3.6:35b-a3b", input=history, tools=tools)
# Stateless: add the model's output to the history before the tool results.
history += response.output
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
result = {"city": args["city"], "temp_c": 11} # your own tool
history.append({
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
})
response = client.responses.create(model="qwen3.6:35b-a3b", input=history, tools=tools)
print(response.output_text)Supported:
functiontools, including several parallel calls in the same turn (parallel_tool_calls)namespacetools: grouped functions. The response carries bothnameandnamespacecustomtools (freeform input): the response is acustom_tool_callwithinputas a string, and the result is sent ascustom_tool_call_outputtool_choice:auto,none,requiredor a specific function. Unknown values are treated asauto- Images in tool results:
input_imagewith a data URL inoutput
Hosted tools are ignored
OpenAI's hosted tools (web_search, file_search, code_interpreter,
computer_use, image_generation, mcp and others) are removed from the request
without an error. The model can only call function tools that you run yourself.
If you need web search, use the
built-in web search
on chat completions.
More on how the models handle tools in Tool calling.
Streaming
Set stream: true and the response arrives as server-sent events in Responses
format:
stream = client.responses.create(
model="qwen3.6:35b-a3b",
input="Write a haiku about Stockholm.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)| Event | When |
|---|---|
response.created, response.in_progress | As soon as the request is accepted |
response.output_item.added | A new item starts (text, reasoning or tool call) |
response.output_text.delta | Incremental text |
response.reasoning_text.delta | Incremental reasoning (thinking alias) |
response.function_call_arguments.delta | The arguments for a tool call, in one piece |
response.output_text.done, response.function_call_arguments.done, response.output_item.done | An item is finished |
response.completed | The response is finished, with usage |
response.incomplete | max_output_tokens was reached |
response.failed | Error during generation, see Errors |
Every event carries an increasing sequence_number.
Tool calls arrive in one piece
Text and reasoning stream token by token even when the request has tools. A
tool call is sent once the model has finished it: output_item.added, a single
function_call_arguments.delta with all arguments, function_call_arguments.done
and output_item.done arrive back to back. Run the tool once you have
output_item.done.
Other parameters
| Parameter | Behaviour |
|---|---|
model | All chat models, including thinking aliases |
instructions | Prepended as the system prompt. developer messages in input are also treated as system |
input | String or list of message, function_call, function_call_output, custom_tool_call, custom_tool_call_output and reasoning |
max_output_tokens | Maximum generated tokens. When reached, the response gets status: "incomplete" |
temperature, top_p | Passed on to the model |
text.format | json_schema and json_object give structured output |
reasoning.effort | "none" turns thinking off. Other values are ignored: turn thinking on with a thinking alias |
metadata | Echoed back in the response |
store | Ignored, nothing is stored |
reasoning items you send back in input are dropped: earlier reasoning is never
sent to the model again. Images are sent as input_image with a data URL, with the
same models and rules as chat completions.
Usage
usage in response.completed (and in non-streamed responses) uses the Responses
format: input_tokens, input_tokens_details.cached_tokens (prompt cache hits),
output_tokens, output_tokens_details.reasoning_tokens and total_tokens.
Billing is the same as for chat completions.
Errors
Errors are returned in OpenAI format:
{"error": {"message", "type", "param", "code"}}.
| Status | When |
|---|---|
| 400 | previous_response_id, conversation, background, item_reference, input_file or an image via file_id (all need stored state), or an invalid body. param names the field |
400 context_length_exceeded | The request doesn't fit the model's context window |
401 invalid_api_key | Missing or invalid key |
| 429 | Token limit (insufficient_quota) or capacity ceiling (capacity_exceeded), with Retry-After |
If the error happens mid-stream it arrives as response.failed with code
context_length_exceeded, invalid_prompt or server_error (transient, retry).
See Errors for the full list.
Codex CLI
Codex speaks the Responses API and works against this endpoint. The configuration is under Codex CLI.