Skip to content
Open benchmarks for local AI

What can your hardware do?

The model, the precision, the workload and the container are fixed. Only your hardware changes. The numbers compare.

For hardware that you can run at home. LocalMax may accept a result from a data-centre GPU or system, but it reserves the right not to show it and not to rank it.

Read the code first

Source on GitHub

LocalMax runs on your machine. Examine the source before you run it. Every part is public: the runner, the container files, the profiles and the checks.

Collected

GPU model, VRAM, driver, CUDA, PCIe link, power limit and clocks. CPU model and core count. RAM. OS and kernel. All measured values, the raw records for each request, and the telemetry.

Never collected

Hostname. Username. File paths. GPU serial numbers. Board UUIDs. MAC addresses. Environment variables. Data about other processes. The runner removes secrets and paths from logs on your machine, before it writes them.

Your name is generated

LocalMax gives your machine a name such as Rocinante-K7X2P. The name comes from a key on your machine. You do not type a name. Keep the 5-character code. Use it to find your results at /systems/CODE.

Submission is voluntary. You control what you send. Commandlocalmax inspect shows the full content before upload. Legal · Privacy

Results
0
— ranked
Systems
0
GPU models
0
Profiles
15
3 categories, 3 tiers, 3 lanes

Three tiers

One model cannot span 12 GB to 192 GB. A tier is the VRAM its run must fill. Each tier has its own model.

Entry · 12 GB+
up to 4B parameters at FP8
The smallest supported card runs this.
0 results
Enthusiast · 24 GB+
up to 8B parameters at FP8
Sized to fill 24 GB.
0 results
Prospector · 64 GB+
up to 32B parameters at FP8
Sized to fill 64 GB and more. A cluster counts as one system.
0 results

A tier gives the largest model it runs, in billions of parameters, at FP8. LocalMax also accepts and ranks INT4 and NVFP4 submissions. Each lane has its own leaderboard, because different precisions are different work.Methodology →

Latest results

All results →
No results yet. The Entry profile runs on a 12 GB card in about 14 minutes.
Fixed
Model, revision, precision, runtime, flags, prompts, token counts, images, steps and seeds. The profile holds all of them. LocalMax checks them at submission.
Measured
Decode and prefill throughput, time to first token, inter-token latency, seconds per step, images per minute, peak VRAM, power, energy and throttle events.
Published
The manifest, the raw records, the telemetry and the system report. LocalMax computes each ranked number again from the raw data before it accepts a result.