Skip to content
Benchmarks / LLM /Prospector / NVFP4

LLM · Prospector · NVFP4

A 72B model at NVFP4, sized to fill 64 GB. NVFP4 needs Blackwell.

Best Decode
tok/s
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run

Best result per GPU

Decode · tok/s
No results yet for this view.
Show the values in this chart (0 rows)
GPUGPUsSamplesBest (tok/s)Mean (tok/s)

Download as CSV· Full open dataset

No results match these filters.

What is pinned

The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.

Model
Repository
nvidia/Qwen2.5-72B-Instruct-NVFP4
Revision
pending freeze
Precision
nvfp4
Parameters
72 B
Licence
Qwen
Runtime
Engine
vllm 0.26.0
Harness
aiperf
dtype
bfloat16
max-model-len
32768
gpu-memory-utilization
0.9
max-num-seqs
16
enforce-eager
false
disable-log-requests
true
tensor-parallel-size
1
swap-space
0
Workloads
interactive
input_tokens=1024 output_tokens=256 concurrency=1 input_source=fixed_synthetic input_seed=20260801
concurrency
input_tokens=1024 output_tokens=256 concurrency_sweep=1/2/4/8 input_source=fixed_synthetic input_seed=20260801
longcontext
input_tokens=8192 output_tokens=1024 concurrency=1 input_source=fixed_synthetic input_seed=20260801
Ranking
Ranked on
Decode (tok/s, higher is better)
Gate
TTFT p95 ≤ 3000
Gate
Inter-token p95 ≤ 200
Gate
Error rate ≤ 0
Also shown
Prefill, TTFT p50, TTFT p95, Inter-token p95, Peak throughput @c8, Energy / token, Model load
Design notes
  • This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
  • A 32B model at FP8 fits a DGX Spark and a dual RTX PRO 6000 workstation. A 70B model would exclude the Spark from its own tier.
  • The RTX PRO 6000 Blackwell has no NVLink. A 2-GPU result uses PCIe. The result page shows the PCIe generation and width.
  • This tier accepts a cluster. LocalMax computes the tier from the total VRAM across all nodes, so eight DGX Sparks are one Prospector system. It ranks each GPU count and each parallelism mode separately, because the network fabric changes the result.