API reference
The ReadyServer LLM API is served by vLLM and follows the OpenAI API format, so most OpenAI clients work once you change the base URL.
Base URL
http://api.readyserver.ai/v1
Reachable from every ReadyServer VPS over our private network, and not from the public internet. Paths below are relative to this URL.
Authentication
None. Access is controlled by the network: only ReadyServer VPS can reach the API. If your client insists on an API key, send any non-empty value, such as not-needed.
Models
| Model ID | Context length | Thinking |
|---|---|---|
qwen3.6-27b | 131,072 tokens | Off: answers directly |
qwen3.8-27b | 262,144 tokens | On by default (adjustable) |
Both models accept text and images, and support tool calling and structured output. See Models for details.
Endpoints
| Method | Path | Purpose | Status |
|---|---|---|---|
| GET | /models | List the models currently served | Supported |
| POST | /chat/completions | Chat, with streaming, tools, structured output and images | Supported |
| POST | /completions | Text completions (legacy format) | Supported |
| POST | /messages | Anthropic Messages API format | Supported |
| POST | /responses | OpenAI Responses API | Preview |
| POST | /embeddings | Text embeddings | Not available |
The Responses API is in preview. Use Chat Completions for production workloads.
Common parameters
For POST /chat/completions:
| Parameter | Notes |
|---|---|
model | Required. qwen3.6-27b or qwen3.8-27b. |
messages | Required. The conversation as a list of {"role", "content"} objects. |
max_tokens | Maximum tokens to generate. If you leave it out, generation can run up to the model's context length, so set it. |
temperature, top_p | Sampling controls. See Models for recommended values. |
stream | true to receive tokens as server-sent events while they're generated. Add "stream_options": {"include_usage": true} to get token usage in the final chunk. |
stop | One or more strings that end generation. |
tools, tool_choice | Tool calling. Supported on both models. |
response_format | Structured output: json_object or json_schema. |
reasoning_effort | qwen3.8-27b only: "low" or "medium". See Thinking. |
logprobs, top_logprobs | Return log probabilities for generated tokens. |
vLLM accepts further parameters, such as top_k and chat_template_kwargs; see the vLLM documentation. With the OpenAI SDKs, pass these through extra_body.
Thinking
qwen3.8-27b thinks before it answers. Its reasoning is returned separately from the answer, in the reasoning field of the message (or of each delta when streaming), and counts towards completion_tokens.
Thinking helps with multi-step problems but takes longer and uses more tokens. By default the model uses its highest reasoning effort. For simpler tasks, lower the effort or turn thinking off:
Python
from openai import OpenAI
client = OpenAI()
# Less thinking: reasoning_effort can be "low" or "medium"
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Is 221 a prime number?"}],
reasoning_effort="low",
max_tokens=1000,
)
print(response.choices[0].message.reasoning)
print(response.choices[0].message.content)
# No thinking: the model answers directly
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Is 221 a prime number?"}],
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
max_tokens=200,
)
print(response.choices[0].message.content)
curl
curl "$OPENAI_BASE_URL/chat/completions" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Is 221 a prime number?"}],
"chat_template_kwargs": {"enable_thinking": false},
"max_tokens": 200
}'
When thinking is on, the answer can start with blank lines; trim them if they matter to you. qwen3.6-27b answers directly, without a thinking step.
Tool calling
Describe your functions in tools. When the model wants to call one, the response has finish_reason set to "tool_calls" and the calls in message.tool_calls:
import json
from openai import OpenAI
client = OpenAI()
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
response = client.chat.completions.create(
model="qwen3.6-27b",
messages=[{"role": "user", "content": "What's the weather in Singapore?"}],
tools=tools,
max_tokens=300,
)
for call in response.choices[0].message.tool_calls or []:
print(call.function.name, json.loads(call.function.arguments))
Run the function yourself, then send its result back as a message with "role": "tool" and the matching tool_call_id so the model can write its answer.
Structured output
Use response_format to get JSON that matches a schema:
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="qwen3.6-27b",
messages=[{"role": "user", "content": "Give the capital of Japan."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "capital",
"schema": {
"type": "object",
"properties": {
"country": {"type": "string"},
"capital": {"type": "string"},
},
"required": ["country", "capital"],
},
},
},
max_tokens=100,
)
print(response.choices[0].message.content) # a JSON object with country and capital
For JSON without a fixed schema, use {"type": "json_object"} and describe the shape you want in the prompt.
Image input
Both models accept images. Send them as base64 data URLs alongside your text:
import base64
from openai import OpenAI
client = OpenAI()
with open("image.png", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="qwen3.6-27b",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in one sentence."},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
],
}],
max_tokens=200,
)
print(response.choices[0].message.content)
Images count towards the prompt's tokens and the context length.
Anthropic Messages API
Tools and SDKs built for the Anthropic Messages API can use the same models. Set the base URL to http://api.readyserver.ai (without /v1) and use any placeholder API key:
# pip install anthropic
import anthropic
client = anthropic.Anthropic(base_url="http://api.readyserver.ai", api_key="not-needed")
message = client.messages.create(
model="qwen3.6-27b",
max_tokens=200,
messages=[{"role": "user", "content": "Explain what a VPS is in one sentence."}],
)
print(message.content[0].text)
Limits
| Limit | Value |
|---|---|
| Context length (prompt + output) | 131,072 tokens (qwen3.6-27b)262,144 tokens ( qwen3.8-27b) |
| Output tokens per request | Up to the context length, minus the prompt |
When the service is busy, requests are queued and customers who have used fewer tokens are served first. See the fair use policy.
Errors
| Status | Meaning | What to do |
|---|---|---|
400 | Invalid request, for example max_tokens or the prompt exceeds the context length | Fix the request; the error message says what's wrong |
404 | Model or endpoint not found | Use an id from GET /models and check the path |
500, 502, 503 | Server error, or the service is temporarily unavailable | Retry with backoff |
Error responses use the OpenAI format: {"error": {"message": "...", "type": "...", "code": 400}}.
Retries and timeouts
- Retry
5xxresponses with exponential backoff and jitter (for example 1 s, 2 s, 4 s, up to a cap). If the response includes aRetry-Afterheader, wait at least that long. - The official OpenAI SDKs already retry these errors a few times by default. Adjust with the
max_retriesclient option. - Don't resend a
400request unchanged; it will fail again. - When the service is busy, requests wait in the queue before generation starts. Use a generous client timeout and stream responses, rather than retrying a request that is only slow: retries add to the queue.
Further reading
- OpenAI API reference: request and response formats.
- vLLM documentation: the inference engine behind the API.