$npx skillfedfor your agent

genai-perf

GenAI Perf Analyzer CLI - CLI tool to simplify profiling LLMs and Generative AI models with Perf Analyzer

With conditionsPyPI Software DevelopmentReleased Aug 2025314.5K downloads / moBSDPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — genai_perf-0.0.16-py3-none-any.whl
v0.0.16 · released 2025-08-26 · Python <4,>=3.10 · 19 runtime deps: fastparquet, jinja2, kaleido, numpy, optuna, orjson, pandas, perf-analyzer

Yes, if you are benchmarking generative AI models on Triton Inference Server or compatible inference servers and need detailed token-level and request-level metrics. The low install friction and active maintenance make it a practical choice. Requires CUDA 12 and a running inference server; the large dependency tree (19 runtime packages) may add setup time. Not suitable if you need to profile models without an external inference server or on non-Triton platforms.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires CUDA 12 to be pre-installed on the system and an inference server (e.g., Triton) already running at the specified endpoint.
  • Low install friction with a pure Python wheel.
  • Requires Python 3.10 or 3.12 and CUDA 12 to be pre-installed; the 19 runtime dependencies include heavy data science and ML stacks (transformers, pandas, numpy, plotly, statsmodels) which may take time to resolve.

License · maintenance · safety

BSD (permissive) — BSD permissive license allows commercial and private use with minimal restrictions; you must retain copyright notices and disclaimers in redistributions.

last release 2025-08-26 (353 days) · last repo commit 2026-08-07 · 153 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 314,511 downloads/mo, #7,698 on PyPI

Verify before relying

pip install genai-perf
genai-perf profile -m gpt2 --backend tensorrtllm --streaming
  • Whether the tool works with inference servers other than Triton Inference Server.
  • Performance overhead of the profiling tool itself on measured metrics.
  • Compatibility with custom model backends beyond those documented in the description.
Same gist for agents: .md · .json

What it is and what it does

GenAI-Perf is a profiling and benchmarking tool designed to measure the performance characteristics of generative AI models running on inference servers. It generates configurable load (concurrent requests or request rates) against a running inference server and collects detailed metrics including output token throughput, time to first token, inter-token latency, and request latency. Results are reported in console tables and exported to CSV and JSON for further analysis.

The tool targets a wide range of model types—large language models, multi-modal models, embeddings, ranking models, and LoRA-adapted variants—and supports both synthetic load generation and real input datasets. It can be configured via command-line arguments or YAML configuration files, and provides customizable frontends and Jinja2-templated payloads for benchmarking custom APIs. The package is in active development (Alpha status) and requires an external inference server to already be running.

Use it for

  • Benchmark LLM inference latency and throughput on Triton Inference Server with TensorRT-LLM backends.
  • Measure time-to-first-token and inter-token latency for streaming language model deployments.
  • Profile multi-modal model performance under concurrent request loads to identify bottlenecks.
  • Compare inference performance across different model backends or hardware configurations.
  • Generate performance reports (CSV/JSON) for embedding or ranking models to track optimization progress.
  • Test custom API endpoints with templated payloads to validate inference server integration.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are benchmarking generative AI models on Triton Inference Server or compatible inference servers and need detailed token-level and request-level metrics.

The low install friction and active maintenance make it a practical choice. Requires CUDA 12 and a running inference server; the large dependency tree (19 runtime packages) may add setup time. Not suitable if you need to profile models without an external inference server or on non-Triton platforms.

Install

genai-perf on PyPI

Before you install

Low install friction with a pure Python wheel. Requires Python 3.10 or 3.12 and CUDA 12 to be pre-installed; the 19 runtime dependencies include heavy data science and ML stacks (transformers, pandas, numpy, plotly, statsmodels) which may take time to resolve. Actively maintained with recent commits.

Requires CUDA 12 to be pre-installed on the system and an inference server (e.g., Triton) already running at the specified endpoint.

License in practice

BSD permissive license allows commercial and private use with minimal restrictions; you must retain copyright notices and disclaimers in redistributions.

Quickstart

pip install genai-perf
genai-perf profile -m gpt2 --backend tensorrtllm --streaming

Verify before relying

  • Whether the tool works with inference servers other than Triton Inference Server.
  • Performance overhead of the profiling tool itself on measured metrics.
  • Compatibility with custom model backends beyond those documented in the description.

Package facts

LicenseBSD permissive
Python supportSupports the current Python release <4,>=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
19 packages
fastparquetjinja2kaleidonumpyoptunaorjsonpandasperf-analyzerpillowplotlypyarrowpytestpytest-mockpyyamlresponsesrichsoundfilestatsmodelstransformers
MaintenanceActively maintained 353 days since the last release
Last repo commit
First released
Downloads314,511 / month, #7,698 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: UnixProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.12Topic :: Scientific/EngineeringTopic :: Software Development

Evidence: genai_perf-0.0.16-py3-none-any.whl

Tags

Capabilities
llm performance benchmarkinginference server latency measurementtoken throughput profilinggenerative ai model metricsperf analyzer cli tooltime to first token measurementrequest latency benchmarking
Topics
benchmarkinginference-serverllm-profiling

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “llm performance benchmarking”

  • genai-perfGenAI-Perf is a command-line tool for measuring throughput, latency,…
  • unitxtUnitxt provides a unified framework for evaluating AI model…
  • terminal-benchTerminal-Bench provides a benchmark suite and execution harness for…

Give your agent the search over MCP, or paste the wish link into any chat.

More Software Development packages

typing-extensions Worth it
PyPI · Software Development · released Jul 2026

Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.

PSF-2.0pure Python · 3.9+
1.9Bdownloads / mo
numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
annotated-doc With conditions
PyPI · Software Development · released Jul 2026

Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.

MITpure Python · 3.9+
456.2Mdownloads / mo
typer Worth it
PyPI · Software Development · released Aug 2026

Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.

Install it if you are building CLIs in Python.

MITpure Python · 3.10+
369.3Mdownloads / mo
distlib With conditions
PyPI · Software Development · released Jun 2026

Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.

permissive licensepure Python
323.3Mdownloads / mo

See also aiperf · perf-analyzer · gllm-inference-binary · onnxruntime-genai · vllm · tokenspeed-mla · nvidia-modelopt · azure-ai-evaluation · cache-dit · genagent

Further reading