Last updated: 11 September 2026

How fair share works

Every request runs on inference capacity shared by ReadyServer VPS customers. Rather than billing per token, we share that capacity so one heavy workload can't crowd out everyone else.

When the service is busy, requests are queued and customers who have used fewer tokens are served first.

Limits

Each request is limited by the model's context length, which includes the output: 131,072 tokens for qwen3.6-27b and 262,144 tokens for qwen3.8-27b. See Models.

When the service is busy

Requests wait in the queue, so responses can take longer to start. Customers who have used fewer tokens are served first.

  • Use a client timeout long enough to wait in the queue, and stream responses.
  • Don't retry requests that are only slow: retries add to the queue.
  • If you get a 5xx error, retry with exponential backoff. See Retries and timeouts.

Acceptable use

  • Use the API for workloads running on your own ReadyServer VPS.
  • Don't resell access to the API or expose the endpoint to the public internet, for example through a public reverse proxy.
  • Don't use the API to create illegal content, spam, malware or material that harasses others.
  • Don't run load tests or benchmarks designed to saturate the service.
  • Follow the licence terms of the models you use.
  • ReadyServer's Terms of Service also apply.

Being a good neighbour

Using fewer tokens also puts you earlier in the queue when the service is busy.

  • Set max_tokens. Requests that stop when they have what they need free up capacity sooner.
  • Use thinking only when it helps. With qwen3.8-27b, lower the reasoning effort or turn thinking off for simple tasks, or use qwen3.6-27b.
  • Stream responses. Your users see output immediately, and you can stop early if the answer is good enough.
  • Limit parallel requests. For batch jobs, process items with a small, fixed concurrency.
  • Cache results. Don't regenerate answers you've already computed.

Enforcement

If a workload affects other customers, we may slow it down or contact you. Repeated or serious misuse may lead to suspension of API access.

Questions

If you have questions about this policy or need more capacity for a project, contact ReadyServer.