Troubleshooting
Fixes for the most common problems when calling the LLM API from your VPS.
Connection times out or is refused
The API is only reachable from a ReadyServer VPS over the private network.
- Make sure you're running the command on your ReadyServer VPS, not on your own computer or another provider's server.
- Check that the base URL exactly matches the one in the quickstart, including any port and path.
- Check that the hostname resolves:
getent hosts api.readyserver.aishould print10.8.8.2. - If your VPS firewall restricts outgoing traffic (for example with
ufworiptables), allow connections to the API address.
The SDK says an API key is missing
The OpenAI SDKs require an API key even though this API doesn't use one. Set any placeholder:
export OPENAI_API_KEY="not-needed"
Or pass it directly, for example OpenAI(api_key="not-needed") in Python.
404 or "model does not exist"
The model value must exactly match an id from the models list:
curl http://api.readyserver.ai/v1/models
Requests take a long time to start
Capacity is shared between all ReadyServer VPS customers. When the service is busy, requests are queued and customers who have used fewer tokens are served first.
- Use a client timeout long enough to wait in the queue, and stream responses.
- Don't retry requests that are only slow: retries add to the queue.
- Use fewer tokens: set
max_tokens, and withqwen3.8-27blower the reasoning effort or turn thinking off. - If you get a
502or503error, retry with exponential backoff.
See the fair use policy for how capacity is shared.
qwen3.8-27b is slow or uses many tokens
qwen3.8-27b thinks before it answers, at its highest reasoning effort by default. For simpler tasks, set reasoning_effort to "low", turn thinking off, or use qwen3.6-27b. See Thinking.
400: the prompt is too long
If the error mentions the maximum context length, your prompt plus max_tokens is larger than the model's context window (see Models). Trim the conversation history, summarise earlier turns, or split long documents into smaller chunks.
The response is cut off
If finish_reason is "length", generation stopped at max_tokens or at the context limit. Raise max_tokens within the limits, or ask the model for a shorter answer.
Responses are slow
Long answers take time to generate, and responses can be slower when the service is busy. Stream responses so the first tokens arrive quickly, and keep prompts and max_tokens no larger than you need.
Still stuck?
Contact ReadyServer support and include:
- Your VPS hostname or service ID
- The exact command or code you ran
- The full error message
- When it happened, including your time zone