{"categories":[{"label":"Software Development","url":"https://skillfed.io/packages/category/software-development/9"},{"label":"Scientific/Engineering","url":"https://skillfed.io/packages/category/scientific-engineering/5"}],"enrichment":{"capability":"GenAI-Perf is a command-line tool for measuring throughput, latency, and token-generation metrics of generative AI models served through an inference server, supporting LLMs, multi-modal models, embeddings, and custom APIs.","skillfed_tags":["benchmarking","inference-server","llm-profiling"],"use_cases":["Benchmark LLM inference latency and throughput on Triton Inference Server with TensorRT-LLM backends.","Measure time-to-first-token and inter-token latency for streaming language model deployments.","Profile multi-modal model performance under concurrent request loads to identify bottlenecks.","Compare inference performance across different model backends or hardware configurations.","Generate performance reports (CSV/JSON) for embedding or ranking models to track optimization progress.","Test custom API endpoints with templated payloads to validate inference server integration."],"what_it_does":"GenAI-Perf is a profiling and benchmarking tool designed to measure the performance characteristics of generative AI models running on inference servers. It generates configurable load (concurrent requests or request rates) against a running inference server and collects detailed metrics including output token throughput, time to first token, inter-token latency, and request latency. Results are reported in console tables and exported to CSV and JSON for further analysis.\n\nThe tool targets a wide range of model types\u2014large language models, multi-modal models, embeddings, ranking models, and LoRA-adapted variants\u2014and supports both synthetic load generation and real input datasets. It can be configured via command-line arguments or YAML configuration files, and provides customizable frontends and Jinja2-templated payloads for benchmarking custom APIs. The package is in active development (Alpha status) and requires an external inference server to already be running.","worth_installing":"Yes, if you are benchmarking generative AI models on Triton Inference Server or compatible inference servers and need detailed token-level and request-level metrics. The low install friction and active maintenance make it a practical choice. Requires CUDA 12 and a running inference server; the large dependency tree (19 runtime packages) may add setup time. Not suitable if you need to profile models without an external inference server or on non-Triton platforms."},"id":"genai-perf","links":{"html":"https://skillfed.io/packages/genai-perf","md":"https://skillfed.io/packages/genai-perf.md","pypi":"https://pypi.org/project/genai-perf/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2025-08-26","license_spdx":null,"license_treatment":"permissive","name":"genai-perf","python_support":"supports_current","summary":"GenAI Perf Analyzer CLI - CLI tool to simplify profiling LLMs and Generative AI models with Perf Analyzer"},"popularity":{"monthly_downloads":314511,"position":7698,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.0.16"}
