Benchmarks / LLM /Enthusiast / INT4
LLM · Enthusiast · INT4
A 32B model at INT4, sized to fill 24 GB. Use this lane to compare an RTX 3090 with newer cards.
Best Decode
—tok/s
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
24 GB
~14 min run
Best result per GPU
Decode · tok/sNo results yet for this view.
Show the values in this chart (0 rows)
| GPU | GPUs | Samples | Best (tok/s) | Mean (tok/s) |
|---|
Leaderboard
Including unranked results →No results match these filters.
What is pinned
The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.
Model
- Repository
- Qwen/Qwen3-32B-AWQ
- Revision
- pending freeze
- Precision
- int4_awq
- Parameters
- 32 B
- Licence
- Apache-2.0
Runtime
- Engine
- vllm 0.26.0
- Harness
- aiperf
- dtype
- bfloat16
- max-model-len
- 16384
- gpu-memory-utilization
- 0.9
- max-num-seqs
- 16
- enforce-eager
- false
- disable-log-requests
- true
- tensor-parallel-size
- 1
- swap-space
- 0
Workloads
- interactive
- input_tokens=1024 output_tokens=256 concurrency=1 input_source=fixed_synthetic input_seed=20260801
- concurrency
- input_tokens=1024 output_tokens=256 concurrency_sweep=1/2/4/8 input_source=fixed_synthetic input_seed=20260801
- longcontext
- input_tokens=8192 output_tokens=1024 concurrency=1 input_source=fixed_synthetic input_seed=20260801
Ranking
- Ranked on
- Decode (tok/s, higher is better)
- Gate
- TTFT p95 ≤ 3000
- Gate
- Inter-token p95 ≤ 200
- Gate
- Error rate ≤ 0
- Also shown
- Prefill, TTFT p50, TTFT p95, Inter-token p95, Peak throughput @c8, Energy / token, Model load
Design notes
- This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
- Decode and prefill are separate values. Decode follows memory bandwidth. Prefill follows compute.
- INT4 has the widest hardware support. It is the only lane that includes Ampere. Use it to compare an RTX 3090 with newer cards.