Benchmarks / Diffusion /Entry / FP8
Diffusion · Entry · FP8
Text to image at 1024 by 1024 on 12 GB. Ranked on seconds for each step. Diffusion is compute-bound, not bandwidth-bound.
Best Seconds / step
—s/step
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
12 GB
~14 min run
Best result per GPU
Seconds / step · s/stepNo results yet for this view.
Show the values in this chart (0 rows)
| GPU | GPUs | Samples | Best (s/step) | Mean (s/step) |
|---|
Leaderboard
Including unranked results →No results match these filters.
What is pinned
The profile fixes all values below. LocalMax checks them at each submission. If the runtime flags do not agree, it publishes the result but does not rank it.
Model
- Repository
- stabilityai/stable-diffusion-xl-base-1.0
- Revision
- pending freeze
- Precision
- fp8_e4m3
- Parameters
- 4 B
- Licence
- CreativeML-Open-RAIL-M++
Runtime
- Engine
- diffusers 0.39.0
- Harness
- diffusion-adapter
- torch_dtype
- float8_e4m3fn
- enable_attention_slicing
- false
- enable_vae_slicing
- false
- enable_model_cpu_offload
- false
- torch_compile
- false
Workloads
- t2i_1024
- width=1024 height=1024 steps=30 scheduler=DPMSolverMultistep guidance_scale=5
Ranking
- Ranked on
- Seconds / step (s/step, lower is better)
- Gate
- Error rate ≤ 0
- Gate
- CPU offload ≤ 0
- Also shown
- Images / min, Per image p50, Per image p95, Energy / image, Pipeline load
Design notes
- This profile is a release candidate. The runtime is pinned. The model revision is not final. This leaderboard is provisional.
- Ranked on seconds for each step. That value does not change if the resolution or the step count changes. Images for each minute comes from it.
- Diffusion is compute-bound. The LLM profiles are bandwidth-bound. This is deliberate.
- CPU offload is off, and a gate enforces it. A run that offloads measures the host as well as the GPU. LocalMax publishes it but does not rank it.