aiperf
AIPerf is a package for performance testing of AI models
Decision gist · record as of 2026-08-14
Yes. AIPerf is actively maintained, has no known vulnerabilities, and provides a comprehensive benchmarking solution for generative AI inference endpoints. Install friction is low on common platforms (wheels available for x86_64 and macOS); Linux aarch64 users need a C compiler. The permissive Apache-2.0 license poses no barrier. Install it if you need to measure or compare LLM inference performance across endpoints, models, or load patterns.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- On Linux aarch64, install build-essential (Debian/Ubuntu) or Development Tools (RHEL/CentOS) before pip install aiperf, as the crick dependency requires a C compiler.
- Low install friction with pre-built wheels on Linux x86_64, macOS, and Windows.
- On Linux aarch64, the crick dependency requires a C compiler and build toolchain at install time.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions; you must include a copy of the license and note any modifications.
last release 2026-08-06 (8 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 394,283 downloads/mo, #6,994 on PyPI
Alternatives
Verify before relying
pip install aiperf
aiperf profile \
--model "granite4:350m" \
--streaming \
--endpoint-type chat \
--tokenizer ibm-granite/granite-4.0-micro \
--url http://localhost:11434 \
--request-count 10- Whether optional telemetry integrations (MLflow, OpenTelemetry, Weights & Biases) are production-ready or experimental
- Performance overhead of the multiprocess ZMQ architecture on small benchmarks
- Compatibility with inference endpoints beyond OpenAI and NIM APIs
What it is and what it does
AIPerf is a benchmarking tool that loads generative AI inference endpoints with configurable request patterns and collects detailed performance metrics. It measures time-to-first-token, inter-token latency, request latency, token throughput, and sequence lengths, presenting results in real-time dashboards or static reports. The tool uses a multiprocess architecture with 10 services communicating via ZMQ, supporting multiple UI modes (dashboard, simple progress bars, or headless) and benchmarking strategies including concurrency ramping, request-rate control, trace replay, and adaptive scaling.
You provide an endpoint URL, model name, and workload parameters (concurrency, request count, or arrival patterns), and AIPerf drives load against it while collecting metrics. It integrates with public datasets like ShareGPT, supports OpenAI-compatible chat and completion APIs as well as NIM embeddings and rankings, and offers optional telemetry streaming to MLflow, OpenTelemetry, or Weights & Biases. Results export to CSV, JSON, and logs for further analysis.
Use it for
- Profile latency and throughput of a vLLM or TGI deployment under varying concurrency to find performance bottlenecks
- Replay production traffic traces (Bailian, Baseten, BurstGPT, SageMaker) to benchmark how an endpoint handles real-world workloads
- Discover SLA boundaries by running adaptive scaling to find the maximum sustainable concurrency or request rate
- Compare token-level metrics (TTFT, ITL, TPS) across different model sizes or inference frameworks on the same hardware
- Test request cancellation and timeout resilience by configuring timeout thresholds and observing endpoint behavior
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
AIPerf is actively maintained, has no known vulnerabilities, and provides a comprehensive benchmarking solution for generative AI inference endpoints. Install friction is low on common platforms (wheels available for x86_64 and macOS); Linux aarch64 users need a C compiler. The permissive Apache-2.0 license poses no barrier. Install it if you need to measure or compare LLM inference performance across endpoints, models, or load patterns.
Install
aiperf on PyPI
Before you install
Low install friction with pre-built wheels on Linux x86_64, macOS, and Windows. On Linux aarch64, the crick dependency requires a C compiler and build toolchain at install time. Package is actively maintained with a recent release.
On Linux aarch64, install build-essential (Debian/Ubuntu) or Development Tools (RHEL/CentOS) before pip install aiperf, as the crick dependency requires a C compiler.
License in practice
Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions; you must include a copy of the license and note any modifications.
Quickstart
pip install aiperf
aiperf profile \
--model "granite4:350m" \
--streaming \
--endpoint-type chat \
--tokenizer ibm-granite/granite-4.0-micro \
--url http://localhost:11434 \
--request-count 10
Verify before relying
- Whether optional telemetry integrations (MLflow, OpenTelemetry, Weights & Biases) are production-ready or experimental
- Performance overhead of the multiprocess ZMQ architecture on small benchmarks
- Compatibility with inference endpoints beyond OpenAI and NIM APIs
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release <3.14,>=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 45 packagesaiofilesaiohttpcrickcycloptsdash-bootstrap-componentsdashdatasetsfastapiffmpeg-pythonfilelockhuggingface-hubjinja2jmespathkaleidomatplotlibmsgspecnumpynvidia-ml-pyoptunaorjsonpandaspillowplotlyprometheus-clientprotobufpsutilpyarrowpydantic-settingspydanticpyzmq |
| Maintenance | Actively maintained 8 days since the last release |
| First released | |
| Downloads | 394,283 / month, #6,994 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: aiperf-0.12.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ai model inference testing”
- aiperfAIPerf measures the performance of generative AI models served by…
- unitxtUnitxt provides a unified framework for evaluating AI model…
- langchain-fireworksIntegrates Fireworks.ai language models with LangChain, enabling you…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also genai-perf · perf-analyzer · nvidia-nat-eval · nvidia-nat-atif · opentelemetry-semantic-conventions-ai · azure-ai-inference · garak · terminal-bench · nvidia-nat-mcp · deepeval