{"categories":[{"label":"Testing","url":"https://skillfed.io/packages/category/software-development-testing/3"}],"enrichment":{"capability":"AIPerf measures the performance of generative AI models served by inference endpoints, collecting detailed latency, throughput, and token metrics via command-line benchmarking and generating comprehensive reports.","skillfed_tags":["llm-benchmarking","performance-profiling","load-testing"],"use_cases":["Profile latency and throughput of a vLLM or TGI deployment under varying concurrency to find performance bottlenecks","Replay production traffic traces (Bailian, Baseten, BurstGPT, SageMaker) to benchmark how an endpoint handles real-world workloads","Discover SLA boundaries by running adaptive scaling to find the maximum sustainable concurrency or request rate","Compare token-level metrics (TTFT, ITL, TPS) across different model sizes or inference frameworks on the same hardware","Test request cancellation and timeout resilience by configuring timeout thresholds and observing endpoint behavior"],"what_it_does":"AIPerf is a benchmarking tool that loads generative AI inference endpoints with configurable request patterns and collects detailed performance metrics. It measures time-to-first-token, inter-token latency, request latency, token throughput, and sequence lengths, presenting results in real-time dashboards or static reports. The tool uses a multiprocess architecture with 10 services communicating via ZMQ, supporting multiple UI modes (dashboard, simple progress bars, or headless) and benchmarking strategies including concurrency ramping, request-rate control, trace replay, and adaptive scaling.\n\nYou provide an endpoint URL, model name, and workload parameters (concurrency, request count, or arrival patterns), and AIPerf drives load against it while collecting metrics. It integrates with public datasets like ShareGPT, supports OpenAI-compatible chat and completion APIs as well as NIM embeddings and rankings, and offers optional telemetry streaming to MLflow, OpenTelemetry, or Weights & Biases. Results export to CSV, JSON, and logs for further analysis.","worth_installing":"Yes. AIPerf is actively maintained, has no known vulnerabilities, and provides a comprehensive benchmarking solution for generative AI inference endpoints. Install friction is low on common platforms (wheels available for x86_64 and macOS); Linux aarch64 users need a C compiler. The permissive Apache-2.0 license poses no barrier. Install it if you need to measure or compare LLM inference performance across endpoints, models, or load patterns."},"id":"aiperf","links":{"html":"https://skillfed.io/packages/aiperf","md":"https://skillfed.io/packages/aiperf.md","pypi":"https://pypi.org/project/aiperf/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-06","license_spdx":null,"license_treatment":"permissive","name":"aiperf","python_support":"supports_current","summary":"AIPerf is a package for performance testing of AI models"},"popularity":{"monthly_downloads":394283,"position":6994,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.12.0"}
