API reference

The ReadyServer LLM API is served by vLLM and follows the OpenAI API format, so most OpenAI clients work once you change the base URL.

Base URL

http://api.readyserver.ai/v1

Reachable from every ReadyServer VPS over our private network, and not from the public internet. Paths below are relative to this URL.

Authentication

None. Access is controlled by the network: only ReadyServer VPS can reach the API. If your client insists on an API key, send any non-empty value, such as not-needed.

Models

Model IDContext lengthThinking
qwen3.6-27b131,072 tokensOff: answers directly
qwen3.8-27b262,144 tokensOn by default (adjustable)

Both models accept text and images, and support tool calling and structured output. See Models for details.

Endpoints

MethodPathPurposeStatus
GET/modelsList the models currently servedSupported
POST/chat/completionsChat, with streaming, tools, structured output and imagesSupported
POST/completionsText completions (legacy format)Supported
POST/messagesAnthropic Messages API formatSupported
POST/responsesOpenAI Responses APIPreview
POST/embeddingsText embeddingsNot available

The Responses API is in preview. Use Chat Completions for production workloads.

Common parameters

For POST /chat/completions:

ParameterNotes
modelRequired. qwen3.6-27b or qwen3.8-27b.
messagesRequired. The conversation as a list of {"role", "content"} objects.
max_tokensMaximum tokens to generate. If you leave it out, generation can run up to the model's context length, so set it.
temperature, top_pSampling controls. See Models for recommended values.
streamtrue to receive tokens as server-sent events while they're generated. Add "stream_options": {"include_usage": true} to get token usage in the final chunk.
stopOne or more strings that end generation.
tools, tool_choiceTool calling. Supported on both models.
response_formatStructured output: json_object or json_schema.
reasoning_effortqwen3.8-27b only: "low" or "medium". See Thinking.
logprobs, top_logprobsReturn log probabilities for generated tokens.

vLLM accepts further parameters, such as top_k and chat_template_kwargs; see the vLLM documentation. With the OpenAI SDKs, pass these through extra_body.

Thinking

qwen3.8-27b thinks before it answers. Its reasoning is returned separately from the answer, in the reasoning field of the message (or of each delta when streaming), and counts towards completion_tokens.

Thinking helps with multi-step problems but takes longer and uses more tokens. By default the model uses its highest reasoning effort. For simpler tasks, lower the effort or turn thinking off:

Python

from openai import OpenAI

client = OpenAI()

# Less thinking: reasoning_effort can be "low" or "medium"
response = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Is 221 a prime number?"}],
    reasoning_effort="low",
    max_tokens=1000,
)
print(response.choices[0].message.reasoning)
print(response.choices[0].message.content)

# No thinking: the model answers directly
response = client.chat.completions.create(
    model="qwen3.8-27b",
    messages=[{"role": "user", "content": "Is 221 a prime number?"}],
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
    max_tokens=200,
)
print(response.choices[0].message.content)

curl

curl "$OPENAI_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Is 221 a prime number?"}],
    "chat_template_kwargs": {"enable_thinking": false},
    "max_tokens": 200
  }'

When thinking is on, the answer can start with blank lines; trim them if they matter to you. qwen3.6-27b answers directly, without a thinking step.

Tool calling

Describe your functions in tools. When the model wants to call one, the response has finish_reason set to "tool_calls" and the calls in message.tool_calls:

import json
from openai import OpenAI

client = OpenAI()

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

response = client.chat.completions.create(
    model="qwen3.6-27b",
    messages=[{"role": "user", "content": "What's the weather in Singapore?"}],
    tools=tools,
    max_tokens=300,
)
for call in response.choices[0].message.tool_calls or []:
    print(call.function.name, json.loads(call.function.arguments))

Run the function yourself, then send its result back as a message with "role": "tool" and the matching tool_call_id so the model can write its answer.

Structured output

Use response_format to get JSON that matches a schema:

from openai import OpenAI

client = OpenAI()

response = client.chat.completions.create(
    model="qwen3.6-27b",
    messages=[{"role": "user", "content": "Give the capital of Japan."}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "capital",
            "schema": {
                "type": "object",
                "properties": {
                    "country": {"type": "string"},
                    "capital": {"type": "string"},
                },
                "required": ["country", "capital"],
            },
        },
    },
    max_tokens=100,
)
print(response.choices[0].message.content)  # a JSON object with country and capital

For JSON without a fixed schema, use {"type": "json_object"} and describe the shape you want in the prompt.

Image input

Both models accept images. Send them as base64 data URLs alongside your text:

import base64
from openai import OpenAI

client = OpenAI()

with open("image.png", "rb") as f:
    image = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="qwen3.6-27b",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image in one sentence."},
            {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
        ],
    }],
    max_tokens=200,
)
print(response.choices[0].message.content)

Images count towards the prompt's tokens and the context length.

Anthropic Messages API

Tools and SDKs built for the Anthropic Messages API can use the same models. Set the base URL to http://api.readyserver.ai (without /v1) and use any placeholder API key:

# pip install anthropic
import anthropic

client = anthropic.Anthropic(base_url="http://api.readyserver.ai", api_key="not-needed")

message = client.messages.create(
    model="qwen3.6-27b",
    max_tokens=200,
    messages=[{"role": "user", "content": "Explain what a VPS is in one sentence."}],
)
print(message.content[0].text)

Limits

LimitValue
Context length (prompt + output)131,072 tokens (qwen3.6-27b)
262,144 tokens (qwen3.8-27b)
Output tokens per requestUp to the context length, minus the prompt

When the service is busy, requests are queued and customers who have used fewer tokens are served first. See the fair use policy.

Errors

StatusMeaningWhat to do
400Invalid request, for example max_tokens or the prompt exceeds the context lengthFix the request; the error message says what's wrong
404Model or endpoint not foundUse an id from GET /models and check the path
500, 502, 503Server error, or the service is temporarily unavailableRetry with backoff

Error responses use the OpenAI format: {"error": {"message": "...", "type": "...", "code": 400}}.

Retries and timeouts

Further reading