Know exactly how your AI performs under load.
InferGauge benchmarks LLM endpoints with live dashboards, reproducible benchmarks, and actionable performance insights. Measure latency, TTFT, throughput, goodput, cost, and capacity from your own machine.
Getting started
One command to install.
curl -fsSL https://raw.githubusercontent.com/Nexus-InferGauge/infergauge-releases/main/install.sh | shThree commands to your first result.
Capabilities
Everything you need to benchmark AI inference.
Concurrency
Find the point where your application starts to slow down as user count increases.
Latency
Measure average, p95, and p99 latency under realistic workloads.
Time to First Token
Understand queueing delays before responses begin streaming.
Goodput
Track the percentage of requests that actually meet your SLOs.
Cost
Estimate token usage, cost per request, and projected spend.
Performance Score
Summarize every run with transparent metrics and reproducible results.
Built for developers.
Runs locally
Prompts, API keys, and benchmark results stay on your machine.
Reproducible
Every benchmark is a small YAML file that anyone on your team can rerun.
CI ready
Detect regressions, compare runs, and automate performance testing.
Simple workflow.
Configure your endpoint
Define your benchmark with a small YAML configuration.
Run a benchmark
Launch tests from the CLI against local or hosted models.
Understand the results
Explore live dashboards, compare runs, and export reports.
Start measuring without guesswork.
Benchmark your AI applications, understand their limits, and ship with confidence.