Responses (OpenAI)

Last updated:

On this page
text
POST https://api.staik.se/v1/responses

Endpoint compatible with OpenAI's Responses API. The OpenAI SDKs' responses.create, Codex CLI and other clients that speak the Responses protocol work directly. Authenticate with Authorization: Bearer (same key).

Requests run on the same path as chat completions: same models, quotas, queues, failover and billing. Usage is recorded under the endpoint responses.

python
from openai import OpenAI

client = OpenAI(api_key="sk-st-your-key", base_url="https://api.staik.se/v1")

response = client.responses.create(
    model="qwen3.6:35b-a3b",
    instructions="Answer briefly.",
    input="Explain how a transformer works.",
)
print(response.output_text)

Stateless: send the full history

staik never stores responses. Every request must contain the full conversation in input, including earlier model output, tool calls and tool results. store is ignored and the response always has "store": false.

Function tools

Tools use the Responses format (type: "function" with name and parameters directly on the tool). The model answers with function_call items, you run the tool and send the result back as function_call_output with the same call_id:

python
import json

tools = [{
    "type": "function",
    "name": "get_weather",
    "description": "Get the current weather for a city.",
    "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
}]

history = [{"role": "user", "content": "What's the weather in Gothenburg?"}]
response = client.responses.create(model="qwen3.6:35b-a3b", input=history, tools=tools)

# Stateless: add the model's output to the history before the tool results.
history += response.output
for item in response.output:
    if item.type == "function_call":
        args = json.loads(item.arguments)
        result = {"city": args["city"], "temp_c": 11}  # your own tool
        history.append({
            "type": "function_call_output",
            "call_id": item.call_id,
            "output": json.dumps(result),
        })

response = client.responses.create(model="qwen3.6:35b-a3b", input=history, tools=tools)
print(response.output_text)

Supported:

  • function tools, including several parallel calls in the same turn (parallel_tool_calls)
  • namespace tools: grouped functions. The response carries both name and namespace
  • custom tools (freeform input): the response is a custom_tool_call with input as a string, and the result is sent as custom_tool_call_output
  • tool_choice: auto, none, required or a specific function. Unknown values are treated as auto
  • Images in tool results: input_image with a data URL in output

Hosted tools are ignored

OpenAI's hosted tools (web_search, file_search, code_interpreter, computer_use, image_generation, mcp and others) are removed from the request without an error. The model can only call function tools that you run yourself. If you need web search, use the built-in web search on chat completions.

More on how the models handle tools in Tool calling.

Streaming

Set stream: true and the response arrives as server-sent events in Responses format:

python
stream = client.responses.create(
    model="qwen3.6:35b-a3b",
    input="Write a haiku about Stockholm.",
    stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
EventWhen
response.created, response.in_progressAs soon as the request is accepted
response.output_item.addedA new item starts (text, reasoning or tool call)
response.output_text.deltaIncremental text
response.reasoning_text.deltaIncremental reasoning (thinking alias)
response.function_call_arguments.deltaThe arguments for a tool call, in one piece
response.output_text.done, response.function_call_arguments.done, response.output_item.doneAn item is finished
response.completedThe response is finished, with usage
response.incompletemax_output_tokens was reached
response.failedError during generation, see Errors

Every event carries an increasing sequence_number.

Tool calls arrive in one piece

Text and reasoning stream token by token even when the request has tools. A tool call is sent once the model has finished it: output_item.added, a single function_call_arguments.delta with all arguments, function_call_arguments.done and output_item.done arrive back to back. Run the tool once you have output_item.done.

Other parameters

ParameterBehaviour
modelAll chat models, including thinking aliases
instructionsPrepended as the system prompt. developer messages in input are also treated as system
inputString or list of message, function_call, function_call_output, custom_tool_call, custom_tool_call_output and reasoning
max_output_tokensMaximum generated tokens. When reached, the response gets status: "incomplete"
temperature, top_pPassed on to the model
text.formatjson_schema and json_object give structured output
reasoning.effort"none" turns thinking off. Other values are ignored: turn thinking on with a thinking alias
metadataEchoed back in the response
storeIgnored, nothing is stored

reasoning items you send back in input are dropped: earlier reasoning is never sent to the model again. Images are sent as input_image with a data URL, with the same models and rules as chat completions.

Usage

usage in response.completed (and in non-streamed responses) uses the Responses format: input_tokens, input_tokens_details.cached_tokens (prompt cache hits), output_tokens, output_tokens_details.reasoning_tokens and total_tokens. Billing is the same as for chat completions.

Errors

Errors are returned in OpenAI format: {"error": {"message", "type", "param", "code"}}.

StatusWhen
400previous_response_id, conversation, background, item_reference, input_file or an image via file_id (all need stored state), or an invalid body. param names the field
400 context_length_exceededThe request doesn't fit the model's context window
401 invalid_api_keyMissing or invalid key
429Token limit (insufficient_quota) or capacity ceiling (capacity_exceeded), with Retry-After

If the error happens mid-stream it arrives as response.failed with code context_length_exceeded, invalid_prompt or server_error (transient, retry). See Errors for the full list.

Codex CLI

Codex speaks the Responses API and works against this endpoint. The configuration is under Codex CLI.