benchmarksqwenperformance
Qwen 3.5 397B Performance Benchmarks on NVIDIA B200
Engineering Team
Benchmark Setup
We benchmarked Qwen 3.5 397B on our NVIDIA B200 GPU cluster under production-like conditions. Here is what we measured.
Hardware Configuration
- GPU: NVIDIA B200 (Blackwell architecture, 192GB HBM3e)
- Precision: FP8 with FlashAttention-3
- Batch sizes: 1, 4, 8, 16, 32
- Sequence lengths: 512, 2048, 8192, 32768 tokens
Throughput Results
| Batch Size | 512 tokens | 2048 tokens | 8192 tokens | 32768 tokens |
|---|---|---|---|---|
| 1 | 45 tok/s | 38 tok/s | 28 tok/s | 18 tok/s |
| 4 | 160 tok/s | 140 tok/s | 100 tok/s | 65 tok/s |
| 8 | 290 tok/s | 250 tok/s | 180 tok/s | 110 tok/s |
| 16 | 480 tok/s | 400 tok/s | 290 tok/s | 175 tok/s |
| 32 | 720 tok/s | 580 tok/s | 400 tok/s | 240 tok/s |
Streaming Latency
Time-to-First-Token (TTFT) for a 2048-token prompt:
- Batch size 1: 85ms
- Batch size 4: 120ms
- Batch size 8: 180ms
All measurements taken at steady state with continuous request flow.
Key Takeaways
- Linear scaling up to batch size 8 with diminishing returns beyond 16
- Sub-100ms TTFT at batch size 1 — suitable for real-time chat applications
- FP8 precision delivers 1.8x throughput improvement over FP16 with no measurable quality degradation
- FlashAttention-3 reduces memory footprint by 40% for long-context workloads
Try It Yourself
Run your own benchmarks using our OpenAI-compatible API. The free Pro trial includes 5 hours of inference time — enough to validate these numbers yourself.
