Benchmarks / LLM /Prospector / FP8
LLM · Prospector · FP8
A 32B model at FP8, sized to fill 64 GB. It fits a DGX Spark and a dual RTX PRO 6000 workstation. It accepts 1 or 2 GPUs, ranked separately.
Best Decode
—tok/s
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run
Best result per GPU
Decode · tok/sNo results yet for this view.
Show the values in this chart (0 rows)
| GPU | GPUs | Samples | Best (tok/s) | Mean (tok/s) |
|---|
Leaderboard
Including unranked results →No results match these filters.
What is pinned
The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.
Model
- Repository
- Qwen/Qwen3-32B
- Revision
- pending freeze
- Precision
- fp8_e4m3
- Parameters
- 32 B
- Licence
- Apache-2.0
Runtime
- Engine
- vllm 0.26.0
- Harness
- aiperf
- dtype
- auto
- max-model-len
- 32768
- gpu-memory-utilization
- 0.9
- max-num-seqs
- 16
- enforce-eager
- false
- disable-log-requests
- true
- tensor-parallel-size
- 1
- swap-space
- 0
Workloads
- interactive
- input_tokens=1024 output_tokens=256 concurrency=1 input_source=fixed_synthetic input_seed=20260801
- concurrency
- input_tokens=1024 output_tokens=256 concurrency_sweep=1/2/4/8 input_source=fixed_synthetic input_seed=20260801
- longcontext
- input_tokens=8192 output_tokens=1024 concurrency=1 input_source=fixed_synthetic input_seed=20260801
Ranking
- Ranked on
- Decode (tok/s, higher is better)
- Gate
- TTFT p95 ≤ 3000
- Gate
- Inter-token p95 ≤ 200
- Gate
- Error rate ≤ 0
- Also shown
- Prefill, TTFT p50, TTFT p95, Inter-token p95, Peak throughput @c8, Energy / token, Model load
Design notes
- This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
- A 32B model at FP8 fits a DGX Spark and a dual RTX PRO 6000 workstation. A 70B model would exclude the Spark from its own tier.
- The RTX PRO 6000 Blackwell has no NVLink. A 2-GPU result uses PCIe. The result page shows the PCIe generation and width.
- This tier accepts a cluster. LocalMax computes the tier from the total VRAM across all nodes, so eight DGX Sparks are one Prospector system. It ranks each GPU count and each parallelism mode separately, because the network fabric changes the result.