Speed by product
Qwen3.8-27B runs in MXFP4 on a dedicated server — no other tenants. The numbers below are what you can expect per product.
Starter Ada
Generation~40 ± 20 tok/s
First token (TTFT)2.21 s
Prefill p99~80 ± 20 s
Prefill mean~4 ± 2 s
Starter Blackwell
Generation~50 ± 20 tok/s
First token (TTFT)1.86 s
Prefill p99~80 ± 20 s
Prefill mean~4 ± 2 s
PRO
Generation~60 ± 30 tok/s
First token (TTFT)1.33 s
Prefill p99~80 ± 20 s
Prefill mean~4 ± 2 s
Ultra
Generation~110 ± 40 tok/s
First token (TTFT)0.32 s
Prefill p99~15 ± 10 s
Prefill mean~1 ± 1 s
Ultra is ~6× faster
On average, the Ultra plan delivers roughly 6× the speed of the other plans — that is where the GPU memory bandwidth is the fastest, so both prefill and generation are quickest.
Read the numbers with care
- These speeds are for 1 user at an average 45K-token context window. They can vary a lot depending on how you use the server — bigger contexts, more concurrent users, and longer prompts all change the numbers.
- TTFT also depends on your region vs the server region — the network round-trip can add up to ~4 s, and the gap grows with context size.
- Prefill time scales with prompt length (it is a quadratic feature). The p99 values above reflect long-context agentic workloads; short prompts are much faster.
- All runs: Qwen3.8-27B in MXFP4, single dedicated server, no other tenants. Treat these as typical, not guaranteed.