Quickstart

Make your first request to the ReadyServer LLM API in about five minutes.

Run these commands on your ReadyServer VPS

The API is only reachable over ReadyServer's private network, so it won't respond from your own computer.

Before you start

1. Check you can reach the API

List the models being served:

curl http://api.readyserver.ai/v1/models

You should get a JSON response like this (trimmed):

{
  "object": "list",
  "data": [
    {
      "id": "qwen3.6-27b",
      "object": "model",
      "owned_by": "vllm",
      "max_model_len": 131072
    },
    {
      "id": "qwen3.8-27b",
      "object": "model",
      "owned_by": "vllm",
      "max_model_len": 262144
    }
  ]
}

If the command hangs or the connection is refused, see Troubleshooting.

2. Set environment variables

The OpenAI SDKs and many other tools read these variables, so you only need to set the endpoint once:

export OPENAI_BASE_URL="http://api.readyserver.ai/v1"
export OPENAI_API_KEY="not-needed"

No API key is required. The OpenAI SDKs refuse to start without one, so set any non-empty placeholder. To keep the settings after you log out, add both lines to ~/.bashrc.

3. Send a chat completion

curl

curl "$OPENAI_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.6-27b",
    "messages": [
      {"role": "user", "content": "Explain what a VPS is in one sentence."}
    ],
    "max_tokens": 200
  }'

Python

# pip install openai
from openai import OpenAI

client = OpenAI()  # reads OPENAI_BASE_URL and OPENAI_API_KEY

response = client.chat.completions.create(
    model="qwen3.6-27b",
    messages=[{"role": "user", "content": "Explain what a VPS is in one sentence."}],
    max_tokens=200,
)
print(response.choices[0].message.content)

Node.js

// npm install openai, then save as hello.mjs and run: node hello.mjs
import OpenAI from "openai";

const client = new OpenAI(); // reads OPENAI_BASE_URL and OPENAI_API_KEY

const response = await client.chat.completions.create({
  model: "qwen3.6-27b",
  messages: [{ role: "user", content: "Explain what a VPS is in one sentence." }],
  max_tokens: 200,
});
console.log(response.choices[0].message.content);

4. Stream the response

Streaming sends tokens as they're generated, so users see output straight away. It's the best choice for chat interfaces and long answers.

curl

curl -N "$OPENAI_BASE_URL/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Write a haiku about servers."}],
    "stream": true
  }'

Python

from openai import OpenAI

client = OpenAI()

stream = client.chat.completions.create(
    model="qwen3.6-27b",
    messages=[{"role": "user", "content": "Write a haiku about servers."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

Node.js

import OpenAI from "openai";

const client = new OpenAI();

const stream = await client.chat.completions.create({
  model: "qwen3.6-27b",
  messages: [{ role: "user", content: "Write a haiku about servers." }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");

Use it with other tools

Most tools that support OpenAI-compatible APIs, such as agent frameworks, chat interfaces and coding agents, have a base URL or API base setting. Run the tool on your VPS and set:

SettingValue
Base URLhttp://api.readyserver.ai/v1
Modelqwen3.6-27b or qwen3.8-27b
API keyAny placeholder, for example not-needed

Next steps