Benchmarks / Vision /Prospector / FP8
Vision · Prospector · FP8
A 32B vision model at FP8, sized to fill 64 GB.
Best Images / min
—img/min
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run
Best result per GPU
Images / min · img/minNo results yet for this view.
Show the values in this chart (0 rows)
| GPU | GPUs | Samples | Best (img/min) | Mean (img/min) |
|---|
Leaderboard
Including unranked results →No results match these filters.
What is pinned
The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.
Model
- Repository
- Qwen/Qwen3-VL-32B-Instruct
- Revision
- pending freeze
- Precision
- fp8_e4m3
- Parameters
- 32 B
- Licence
- Apache-2.0
Runtime
- Engine
- vllm 0.26.0
- Harness
- aiperf
- dtype
- auto
- max-model-len
- 32768
- gpu-memory-utilization
- 0.9
- max-num-seqs
- 16
- enforce-eager
- false
- disable-log-requests
- true
- tensor-parallel-size
- 1
- swap-space
- 0
Workloads
- vision_ocr
- task=ocr image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_chart
- task=chart image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_document
- task=document image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_description
- task=description image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_reasoning
- task=reasoning image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_throughput
- task=description image_long_edge_px=1080 max_output_tokens=128 concurrency=2 temperature=0
Ranking
- Ranked on
- Images / min (img/min, higher is better)
- Gate
- Answer accuracy ≥ 90
- Gate
- Error rate ≤ 0
- Also shown
- TTFT p50, TTFT p95, End to end p50, Decode, Image encode p50, Energy / image
Design notes
- This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
- The headline is images for each minute at concurrency 2.
- Quality is a gate, not a score. The OCR, chart and document tasks have fixed expected answers. LocalMax records the description and reasoning tasks but does not score them.
- LocalMax reports image encode time separately from prompt prefill.
- This tier accepts a cluster. LocalMax computes the tier from the total VRAM across all nodes, so eight DGX Sparks are one Prospector system. It ranks each GPU count and each parallelism mode separately, because the network fabric changes the result.