Troubleshooting

Fixes for the most common problems when calling the LLM API from your VPS.

Connection times out or is refused

The API is only reachable from a ReadyServer VPS over the private network.

The SDK says an API key is missing

The OpenAI SDKs require an API key even though this API doesn't use one. Set any placeholder:

export OPENAI_API_KEY="not-needed"

Or pass it directly, for example OpenAI(api_key="not-needed") in Python.

404 or "model does not exist"

The model value must exactly match an id from the models list:

curl http://api.readyserver.ai/v1/models

Requests take a long time to start

Capacity is shared between all ReadyServer VPS customers. When the service is busy, requests are queued and customers who have used fewer tokens are served first.

See the fair use policy for how capacity is shared.

qwen3.8-27b is slow or uses many tokens

qwen3.8-27b thinks before it answers, at its highest reasoning effort by default. For simpler tasks, set reasoning_effort to "low", turn thinking off, or use qwen3.6-27b. See Thinking.

400: the prompt is too long

If the error mentions the maximum context length, your prompt plus max_tokens is larger than the model's context window (see Models). Trim the conversation history, summarise earlier turns, or split long documents into smaller chunks.

The response is cut off

If finish_reason is "length", generation stopped at max_tokens or at the context limit. Raise max_tokens within the limits, or ask the model for a shorter answer.

Responses are slow

Long answers take time to generate, and responses can be slower when the service is busy. Stream responses so the first tokens arrive quickly, and keep prompts and max_tokens no larger than you need.

Still stuck?

Contact ReadyServer support and include: