GInfer · Community benchmarks
G.bench
See what your hardware can do.
A little friendly competition. Compare GInfer performance with the community — same GPU, same model, same workload.
Community results
Loading community results…
Community leaderboard
The first score could be yours.
Once submissions open, run the Standard Benchmark in GChat and share your result here.
Explore GChat ↗| Published | Rank | Nickname | Cold PP | PP | Aggregate TG | TG / request | GPU | Model | Engine | TP | Concurrency | OS |
|---|
All rates are tokens per second. — means unavailable, not zero. This board compares the latest 500 submissions. Hardware and model filters default to showing all results. Rank is within the current selection at the selected concurrency; model and weight format are shown for each result. Execution settings, reserved capacities, and full hardware details are available on the run page.
Run ID
Loading run…
Throughput by concurrency
| Concurrency | Cold PP | PP | Aggregate TG | TG / request |
|---|
— means unavailable, not zero. Missing measurements are not connected across gaps in the graph.
Hardware & model
Power, clocks & thermals
| Concurrency | GPU | Power limit (enforced / default) | SM clock MHz (min / median / max) | Memory clock MHz (min / median / max) | Power W (min / median / max) | Temperature °C (min / median / max) | Throttle reasons |
|---|
Sampled about every 100 ms while the measured round ran, so the range describes that round only. The power limit is the enforced limit at the time of the run. Throttle reasons are those the driver reported during sampling, with the number of samples that showed each. — means the value was not reported.
Run settings
What the scores mean
- Cold PP
- Prompt tokens actually computed per second during the warmup. Reported only when the wave includes a complete uncached prompt. Existing prefix reuse can make this unavailable.
- PP
- Effective prompt throughput during the measured wave, including prefix-cache reuse: all prompt tokens divided by mean time to first token.
- Aggregate TG
- Generation throughput across the active requests, measured during rounds at the selected concurrency. C4 means four simultaneous requests; TG per request is aggregate TG divided by four.
The repeating prompt exercises GInfer’s KV engine, prefix cache, DFlash when enabled, and per-model, per-SM tuned kernels. It intentionally favors prefix reuse and is not a cold-prompt or mixed-production workload.
These results, and your privacy
Community-reported results are not independently verified measurements. A benchmark is one workload, not a promise of everyday performance.
Publishing is opt-in in GChat. The public record contains your nickname, inference-host CPU, RAM, GPU and OS, model configuration, engine version and build details, GPU driver, PCIe link, power limit, sampled clock, power and temperature ranges, throttle reasons, timing counters, and scores. No hostname, account name, GPU serial or UUID, PCI address, model path, credentials, prompt text, or generated text is submitted. WSL CPU and RAM reflect its assigned resources; unavailable hardware measurements stay blank.
The website receives your connection IP. The application stores a keyed hash for rate limiting, not a public IP address. Hosting access logs are subject to the hosting operator’s retention policy. Keep the private submission receipt to delete your result.