Skip to content
Benchmarks

The profile matrix

A profile is the unit of comparison. Two results share a rank only if the category, tier, lane, profile version, runtime, GPU count and parallelism all agree.

Each cell gives the model size in billions of parameters. A tier is the total VRAM its run must fill. For a cluster, LocalMax adds the VRAM of every node, so eight DGX Sparks are one Prospector system.

LLM

Ranked on decode throughput. Latency gates apply.

Vision

Ranked on images for each minute. An accuracy gate applies.
FP8Needs Ada or newer. The default lane.
INT4Needs Ampere or newer. The widest hardware support.
NVFP4Needs Blackwell. RTX 5090, RTX PRO 6000, DGX Spark.
Entry12 GB+
4B parameters
Qwen3-VL-4B-Instruct
12 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined
Enthusiast24 GB+
8B parameters
Qwen3-VL-8B-Instruct
24 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined
Prospector64 GB+
32B parameters
Qwen3-VL-32B-Instruct
64 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined

Diffusion

Ranked on seconds for each step.
FP8Needs Ada or newer. The default lane.
INT4Needs Ampere or newer. The widest hardware support.
NVFP4Needs Blackwell. RTX 5090, RTX PRO 6000, DGX Spark.
Entry12 GB+
4B parameters
stable-diffusion-xl-base-1.0
12 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined
Enthusiast24 GB+
8B parameters
stable-diffusion-3.5-large
24 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined
Prospector64 GB+
12B parameters
FLUX.1-schnell
64 GB minimum · ~14 min
0 resultsAwaiting results
Not defined
Not defined

Why a leaderboard can stay hidden

You can run a profile as soon as it exists. Its leaderboard appears only after two different systems have verified results. One row is not a comparison.

A profile at version 0.x is not frozen. Its model revision is not final. LocalMax ranks a result only against its own profile version.