How fair share works
Every request runs on inference capacity shared by ReadyServer VPS customers. Rather than billing per token, we share that capacity so one heavy workload can't crowd out everyone else.
When the service is busy, requests are queued and customers who have used fewer tokens are served first.
Limits
Each request is limited by the model's context length, which includes the output: 131,072 tokens for qwen3.6-27b and 262,144 tokens for qwen3.8-27b. See Models.
When the service is busy
Requests wait in the queue, so responses can take longer to start. Customers who have used fewer tokens are served first.
- Use a client timeout long enough to wait in the queue, and stream responses.
- Don't retry requests that are only slow: retries add to the queue.
- If you get a
5xxerror, retry with exponential backoff. See Retries and timeouts.
Acceptable use
- Use the API for workloads running on your own ReadyServer VPS.
- Don't resell access to the API or expose the endpoint to the public internet, for example through a public reverse proxy.
- Don't use the API to create illegal content, spam, malware or material that harasses others.
- Don't run load tests or benchmarks designed to saturate the service.
- Follow the licence terms of the models you use.
- ReadyServer's Terms of Service also apply.
Being a good neighbour
Using fewer tokens also puts you earlier in the queue when the service is busy.
- Set
max_tokens. Requests that stop when they have what they need free up capacity sooner. - Use thinking only when it helps. With
qwen3.8-27b, lower the reasoning effort or turn thinking off for simple tasks, or useqwen3.6-27b. - Stream responses. Your users see output immediately, and you can stop early if the answer is good enough.
- Limit parallel requests. For batch jobs, process items with a small, fixed concurrency.
- Cache results. Don't regenerate answers you've already computed.
Enforcement
If a workload affects other customers, we may slow it down or contact you. Repeated or serious misuse may lead to suspension of API access.
Questions
If you have questions about this policy or need more capacity for a project, contact ReadyServer.