Skip to content
Benchmarks / Vision /Prospector / FP8

Vision · Prospector · FP8

A 32B vision model at FP8, sized to fill 64 GB.

Best Images / min
img/min
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run

Best result per GPU

Images / min · img/min
No results yet for this view.
Show the values in this chart (0 rows)
GPUGPUsSamplesBest (img/min)Mean (img/min)

Download as CSV· Full open dataset

No results match these filters.

What is pinned

The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.

Model
Repository
Qwen/Qwen3-VL-32B-Instruct
Revision
pending freeze
Precision
fp8_e4m3
Parameters
32 B
Licence
Apache-2.0
Runtime
Engine
vllm 0.26.0
Harness
aiperf
dtype
auto
max-model-len
32768
gpu-memory-utilization
0.9
max-num-seqs
16
enforce-eager
false
disable-log-requests
true
tensor-parallel-size
1
swap-space
0
Workloads
vision_ocr
task=ocr image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
vision_chart
task=chart image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
vision_document
task=document image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
vision_description
task=description image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
vision_reasoning
task=reasoning image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
vision_throughput
task=description image_long_edge_px=1080 max_output_tokens=128 concurrency=2 temperature=0
Ranking
Ranked on
Images / min (img/min, higher is better)
Gate
Answer accuracy ≥ 90
Gate
Error rate ≤ 0
Also shown
TTFT p50, TTFT p95, End to end p50, Decode, Image encode p50, Energy / image
Design notes
  • This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
  • The headline is images for each minute at concurrency 2.
  • Quality is a gate, not a score. The OCR, chart and document tasks have fixed expected answers. LocalMax records the description and reasoning tasks but does not score them.
  • LocalMax reports image encode time separately from prompt prefill.
  • This tier accepts a cluster. LocalMax computes the tier from the total VRAM across all nodes, so eight DGX Sparks are one Prospector system. It ranks each GPU count and each parallelism mode separately, because the network fabric changes the result.