perf-analyzer
Triton Performance Analyzer
Decision gist · record as of 2026-08-14
Yes, if you are actively optimizing models on Triton Inference Server and need a structured way to measure performance changes. The tool is actively maintained, has no runtime dependencies, and directly addresses the workflow of iterative model tuning. However, verify the unclear license terms first, and note that it requires an external Triton server to be useful—it is not a standalone profiler.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a running Triton Inference Server instance and a model deployed in its repository; the tool is a client that connects to an external server, not a standalone profiler.
- Medium install friction due to platform-specific wheels (manylinux_2_38 for aarch64 and x86_64).
- Active maintenance with recent release (108 days ago), though no runtime dependencies simplifies deployment once installed.
License · maintenance · safety
(unclear) — License treatment is unclear—no SPDX identifier or raw license text provided in metadata. Verify licensing terms before use in proprietary or commercial contexts.
last release 2026-04-28 (108 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 434,444 downloads/mo, #6,689 on PyPI
Alternatives
Verify before relying
pip install perf-analyzer
perf_analyzer -m simple- Python version requirements are unspecified; verify compatibility with your environment.
- Whether genai-perf deprecation notice affects current perf-analyzer maintenance or feature roadmap.
- Exact license terms and any restrictions on commercial use or redistribution.
What it is and what it does
Perf Analyzer is a command-line benchmarking tool for Triton Inference Server that measures model inference performance under realistic load. It supports multiple load modes (concurrency, request rate, custom intervals) and measurement strategies (time windows, count windows) to help developers identify performance bottlenecks and validate optimization changes. The tool works with standard models, sequence models, ensemble models, and decoupled models, and can auto-generate or accept custom input data for testing.
You run it against a live Triton server to collect latency, throughput, and other performance metrics as you adjust model configurations or server settings. It's designed for iterative optimization workflows where you need to measure the impact of each change before moving to the next experiment.
Use it for
- Benchmark latency and throughput of a model before and after applying optimization techniques.
- Profile ensemble or sequence models to identify which component is the performance bottleneck.
- Validate that model changes (quantization, batching, etc.) actually improve end-to-end inference speed.
- Load-test a Triton deployment to find the maximum sustainable request rate or concurrency.
- Compare inference performance across different hardware or Triton configurations.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are actively optimizing models on Triton Inference Server and need a structured way to measure performance changes.
The tool is actively maintained, has no runtime dependencies, and directly addresses the workflow of iterative model tuning. However, verify the unclear license terms first, and note that it requires an external Triton server to be useful—it is not a standalone profiler.
Install
perf-analyzer on PyPI
Before you install
Medium install friction due to platform-specific wheels (manylinux_2_38 for aarch64 and x86_64). Active maintenance with recent release (108 days ago), though no runtime dependencies simplifies deployment once installed.
Requires a running Triton Inference Server instance and a model deployed in its repository; the tool is a client that connects to an external server, not a standalone profiler.
License in practice
License treatment is unclear—no SPDX identifier or raw license text provided in metadata. Verify licensing terms before use in proprietary or commercial contexts.
Quickstart
pip install perf-analyzer
perf_analyzer -m simple
Verify before relying
- Python version requirements are unspecified; verify compatibility with your environment.
- Whether genai-perf deprecation notice affects current perf-analyzer maintenance or feature roadmap.
- Exact license terms and any restrictions on commercial use or redistribution.
Package facts
| License | Not declared unclear |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 108 days since the last release |
| First released | |
| Downloads | 434,444 / month, #6,689 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: perf_analyzer-2.60.0-py3-none-manylinux_2_38_aarch64.whl; perf_analyzer-2.60.0-py3-none-manylinux_2_38_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “triton inference server benchmarking”
- perf-analyzerPerf Analyzer is a CLI tool that measures and optimizes inference…
- genai-perfGenAI-Perf is a command-line tool for measuring throughput, latency,…
- tritonclienttritonclient is a Python client library for communicating with Triton…
Give your agent the search over MCP, or paste the wish link into any chat.
More Monitoring packages
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
Provides generated Python code for OpenTelemetry semantic conventions, enabling standardized attribute naming and constant definitions for instrumentation and telemetry collection.
Install it if you are using OpenTelemetry and want to follow semantic conventions correctly.
Provides the reference implementation of the OpenTelemetry API for collecting and exporting traces, metrics, and logs from Python applications.
Provides the abstract API and interfaces for OpenTelemetry instrumentation in Python, defining how to emit traces, metrics, and logs without tying code to a specific SDK implementation.
Exports OpenTelemetry observability data to an OpenTelemetry Collector using Protobuf-encoded messages over HTTP.
Install it if you are using OpenTelemetry in Python and need to send data to a Collector over HTTP.
Provides automatic instrumentation commands and programmatic APIs to inject distributed tracing into Python applications without code changes, detecting and instrumenting packages used by your program.
Install it if you need distributed tracing without code changes and have compatible instrumented packages in your environment.
See also genai-perf · tritonclient · aiperf · triton-windows · tokenspeed-triton · triton · nvidia-nat-eval · tensorrt-cu12-libs · transformer-engine · tensorrt-cu13