Skip to content
Benchmarks / Diffusion /Enthusiast / FP8

Diffusion · Enthusiast · FP8

Text to image at 1024 by 1024, sized to fill 24 GB.

Best Seconds / step
s/step
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
24 GB
~14 min run

Best result per GPU

Seconds / step · s/step
No results yet for this view.
Show the values in this chart (0 rows)
GPUGPUsSamplesBest (s/step)Mean (s/step)

Download as CSV· Full open dataset

No results match these filters.

What is pinned

The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.

Model
Repository
stabilityai/stable-diffusion-3.5-large
Revision
pending freeze
Precision
fp8_e4m3
Parameters
8 B
Licence
Stability-Community
Runtime
Engine
diffusers 0.39.0
Harness
diffusion-adapter
torch_dtype
float8_e4m3fn
enable_attention_slicing
false
enable_vae_slicing
false
enable_model_cpu_offload
false
torch_compile
false
Workloads
t2i_1024
width=1024 height=1024 steps=30 scheduler=DPMSolverMultistep guidance_scale=5
Ranking
Ranked on
Seconds / step (s/step, lower is better)
Gate
Error rate ≤ 0
Gate
CPU offload ≤ 0
Also shown
Images / min, Per image p50, Per image p95, Energy / image, Pipeline load
Design notes
  • This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
  • Ranked on seconds for each step. That value does not change if the resolution or the step count changes. Images for each minute comes from it.
  • Diffusion is compute-bound. The LLM profiles are bandwidth-bound. This is deliberate.
  • CPU offload is off, and a gate enforces it. A run that offloads measures the host as well as the GPU. LocalMax publishes it but does not rank it.