Benchmarks / Vision /Entry / FP8
Vision · Entry · FP8
A 4B vision model at FP8 on 12 GB. Ranked on images per minute, behind an answer-accuracy gate.
Best Images / min
—img/min
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
12 GB
~14 min run
Best result per GPU
Images / min · img/minNo results yet for this view.
Show the values in this chart (0 rows)
| GPU | GPUs | Samples | Best (img/min) | Mean (img/min) |
|---|
Leaderboard
Including unranked results →No results match these filters.
What is pinned
The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.
Model
- Repository
- Qwen/Qwen3-VL-4B-Instruct
- Revision
- pending freeze
- Precision
- fp8_e4m3
- Parameters
- 4 B
- Licence
- Apache-2.0
Runtime
- Engine
- vllm 0.26.0
- Harness
- aiperf
- dtype
- auto
- max-model-len
- 32768
- gpu-memory-utilization
- 0.9
- max-num-seqs
- 16
- enforce-eager
- false
- disable-log-requests
- true
- tensor-parallel-size
- 1
- swap-space
- 0
Workloads
- vision_ocr
- task=ocr image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_chart
- task=chart image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_document
- task=document image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_description
- task=description image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_reasoning
- task=reasoning image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_throughput
- task=description image_long_edge_px=1080 max_output_tokens=128 concurrency=2 temperature=0
Ranking
- Ranked on
- Images / min (img/min, higher is better)
- Gate
- Answer accuracy ≥ 90
- Gate
- Error rate ≤ 0
- Also shown
- TTFT p50, TTFT p95, End to end p50, Decode, Image encode p50, Energy / image
Design notes
- This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
- The headline is images for each minute at concurrency 2.
- Quality is a gate, not a score. The OCR, chart and document tasks have fixed expected answers. LocalMax records the description and reasoning tasks but does not score them.
- LocalMax reports image encode time separately from prompt prefill.