Which model should I use?

  • qwen3.6-27b answers directly, so responses start quickly and use fewer tokens. Use it when you want fast, direct answers.
  • qwen3.8-27b thinks before it answers and has a longer context window. Use it for multi-step problems, long documents and agents that need to plan.

Qwen3.6-27B-FP8

A 27-billion-parameter Qwen model, quantised to FP8. Answers directly, without a thinking step.

Available

API model ID
qwen3.6-27b
Developer
Qwen team, Alibaba Cloud
Parameters
27B
Precision
FP8
Context length
131,072 tokens
Thinking
Off
Tool calling
Yes
Structured output
Yes: JSON object and JSON schema
Image input
Yes
Recommended settings
temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5
Licence
Apache 2.0
Model card
Qwen/Qwen3.6-27B-FP8

Qwen3.8-27B-FP8

A 27-billion-parameter Qwen model, quantised to FP8, with a longer context window. Thinks before it answers.

Available

API model ID
qwen3.8-27b
Developer
Qwen team, Alibaba Cloud
Parameters
27B
Precision
FP8
Context length
262,144 tokens
Thinking
On by default; lower the effort or turn it off per request
Tool calling
Yes
Structured output
Yes: JSON object and JSON schema
Image input
Yes
Recommended settings
With thinking: temperature 1.0, top_p 0.95, top_k 20
Without thinking: temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5
Licence
Apache 2.0
Model card
Qwen/Qwen3.8-27B-FP8

Recommended settings come from the Qwen model cards. Pass top_k and presence_penalty through extra_body with the OpenAI SDKs.

Check what's being served

To see exactly which models are available right now, run this on your VPS:

curl http://api.readyserver.ai/v1/models

Request a model

We plan to add more models. Once the community forum opens you'll be able to request models there. Until then, contact ReadyServer support.