Which model should I use?
qwen3.6-27b answers directly, so responses start quickly and use fewer tokens. Use it when you want fast, direct answers.
qwen3.8-27b thinks before it answers and has a longer context window. Use it for multi-step problems, long documents and agents that need to plan.
- API model ID
qwen3.6-27b
- Developer
- Qwen team, Alibaba Cloud
- Parameters
- 27B
- Precision
- FP8
- Context length
- 131,072 tokens
- Thinking
- Off
- Tool calling
- Yes
- Structured output
- Yes: JSON object and JSON schema
- Image input
- Yes
- Recommended settings
- temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5
- Licence
- Apache 2.0
- Model card
- Qwen/Qwen3.6-27B-FP8
- API model ID
qwen3.8-27b
- Developer
- Qwen team, Alibaba Cloud
- Parameters
- 27B
- Precision
- FP8
- Context length
- 262,144 tokens
- Thinking
- On by default; lower the effort or turn it off per request
- Tool calling
- Yes
- Structured output
- Yes: JSON object and JSON schema
- Image input
- Yes
- Recommended settings
- With thinking: temperature 1.0, top_p 0.95, top_k 20
Without thinking: temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5
- Licence
- Apache 2.0
- Model card
- Qwen/Qwen3.8-27B-FP8
Recommended settings come from the Qwen model cards. Pass top_k and presence_penalty through extra_body with the OpenAI SDKs.
Check what's being served
To see exactly which models are available right now, run this on your VPS:
curl http://api.readyserver.ai/v1/models
Request a model
We plan to add more models. Once the community forum opens you'll be able to request models there. Until then, contact ReadyServer support.