openvino-genai
Library of the most popular Generative AI model pipelines, optimized execution methods, and samples
Decision gist · record as of 2026-08-14
Yes, if you need to run generative AI models locally and are willing to pre-convert models to OpenVINO IR format. The library significantly reduces boilerplate compared to raw OpenVINO, and active maintenance plus Apache 2.0 licensing make it low-risk. The main gotcha is strict version pinning of openvino_tokenizers; verify compatibility before updating dependencies. Not recommended if you need out-of-the-box support for Hugging Face model formats or if you prefer cloud-based inference.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a model in OpenVINO IR format; version of openvino_tokenizers must match openvino-genai version to avoid ABI incompatibility.
- Medium install friction: prebuilt wheels available for Python 3.10–3.14 across macOS (ARM64), Linux (x86_64 and aarch64), and Windows, but requires version-matched openvino_tokenizers dependency to avoid ABI incompatibility errors.
- Active maintenance with a release 10 days old.
License · maintenance · safety
permissive license (permissive) — Apache License 2.0 permits commercial and private use with minimal restrictions; you must include a copy of the license and note any modifications, but no patent indemnification or trademark restrictions apply to typical use.
last release 2026-08-04 (10 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 115,968 downloads/mo, #12,226 on PyPI
Alternatives
Verify before relying
pip install openvino-genai
import openvino_genai as ov_genai
pipe = ov_genai.LLMPipeline(models_path, "CPU")
print(pipe.generate("The Sun is yellow because", max_new_tokens=100))- Whether GPU acceleration (e.g., Intel Arc, discrete GPU) is supported beyond CPU inference.
- Performance characteristics and memory footprint compared to other local inference frameworks.
- Supported model architectures and whether custom models can be converted from Hugging Face format.
What it is and what it does
openvino-genai is a Python library that wraps OpenVINO's inference runtime to simplify running generative AI models locally. It provides a high-level LLMPipeline class that automatically loads tokenizers, detokenizers, and generation configs from a model directory, reducing boilerplate code to a few lines. The library handles the complexity of the generation process—including beam search, streaming output, and custom generation parameters—while targeting CPU and hardware accelerators through OpenVINO's backend.
The package is designed for developers who want to run language models without managing low-level inference details. It depends on openvino_tokenizers for tokenization and requires models to be pre-converted to OpenVINO IR format. Version alignment between openvino-genai and its dependencies is critical; mismatched versions can cause ABI incompatibility errors. The library supports both Python and C++ interfaces, though the Python API is the primary entry point for most users.
Use it for
- Run a local chat interface or Q&A system using a converted LLM without cloud dependencies.
- Integrate text generation into a Python application with minimal setup overhead.
- Experiment with different beam search and generation parameters on a single machine.
- Deploy generative AI inference on edge devices or servers with Intel hardware acceleration.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to run generative AI models locally and are willing to pre-convert models to OpenVINO IR format.
The library significantly reduces boilerplate compared to raw OpenVINO, and active maintenance plus Apache 2.0 licensing make it low-risk. The main gotcha is strict version pinning of openvino_tokenizers; verify compatibility before updating dependencies. Not recommended if you need out-of-the-box support for Hugging Face model formats or if you prefer cloud-based inference.
Install
openvino-genai on PyPI
Before you install
Medium install friction: prebuilt wheels available for Python 3.10–3.14 across macOS (ARM64), Linux (x86_64 and aarch64), and Windows, but requires version-matched openvino_tokenizers dependency to avoid ABI incompatibility errors. Active maintenance with a release 10 days old.
Requires a model in OpenVINO IR format; version of openvino_tokenizers must match openvino-genai version to avoid ABI incompatibility.
License in practice
Apache License 2.0 permits commercial and private use with minimal restrictions; you must include a copy of the license and note any modifications, but no patent indemnification or trademark restrictions apply to typical use.
Quickstart
pip install openvino-genai
import openvino_genai as ov_genai
pipe = ov_genai.LLMPipeline(models_path, "CPU")
print(pipe.generate("The Sun is yellow because", max_new_tokens=100))
Verify before relying
- Whether GPU acceleration (e.g., Intel Arc, discrete GPU) is supported beyond CPU inference.
- Performance characteristics and memory footprint compared to other local inference frameworks.
- Supported model architectures and whether custom models can be converted from Hugging Face format.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packageopenvino_tokenizers |
| Maintenance | Actively maintained 10 days since the last release |
| First released | |
| Downloads | 115,968 / month, #12,226 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxOperating System :: UnixProgramming Language :: CProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: Implementation :: CPythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development :: Libraries :: Python Modules |
Evidence: openvino_genai-2026.3.0.0-2495-cp310-cp310-macosx_11_0_arm64.whl; openvino_genai-2026.3.0.0-2495-cp310-cp310-manylinux_2_28_x86_64.whl; openvino_genai-2026.3.0.0-2495-cp310-cp310-manylinux_2_31_aarch64.whl; openvino_genai-2026.3.0.0-2495-cp310-cp310-win_amd64.whl; openvino_genai-2026.3.0.0-2495-cp311-cp311-macosx_11_0_arm64.whl; openvino_genai-2026.3.0.0-2495-cp311-cp311-manylinux_2_28_x86_64.whl; openvino_genai-2026.3.0.0-2495-cp311-cp311-manylinux_2_31_aarch64.whl; openvino_genai-2026.3.0.0-2495-cp311-cp311-win_amd64.whl; openvino_genai-2026.3.0.0-2495-cp312-cp312-macosx_11_0_arm64.whl; openvino_genai-2026.3.0.0-2495-cp312-cp312-manylinux_2_28_x86_64.whl; openvino_genai-2026.3.0.0-2495-cp312-cp312-manylinux_2_31_aarch64.whl; openvino_genai-2026.3.0.0-2495-cp312-cp312-win_amd64.whl; openvino_genai-2026.3.0.0-2495-cp313-cp313-macosx_11_0_arm64.whl; openvino_genai-2026.3.0.0-2495-cp313-cp313-manylinux_2_28_x86_64.whl; openvino_genai-2026.3.0.0-2495-cp313-cp313-manylinux_2_31_aarch64.whl; openvino_genai-2026.3.0.0-2495-cp313-cp313-win_amd64.whl; openvino_genai-2026.3.0.0-2495-cp314-cp314-macosx_11_0_arm64.whl; openvino_genai-2026.3.0.0-2495-cp314-cp314-manylinux_2_28_x86_64.whl; openvino_genai-2026.3.0.0-2495-cp314-cp314-manylinux_2_31_aarch64.whl; openvino_genai-2026.3.0.0-2495-cp314-cp314t-macosx_11_0_arm64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “OpenVINO model serving”
- openvino-genaiopenvino-genai simplifies running inference on generative AI models…
- optimum-intelOptimum Intel bridges Hugging Face Transformers and Diffusers models…
- openvino-devProvides command-line tools and Python APIs to convert deep learning…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also openvino · openvino-tokenizers · optimum-intel · openvino-dev · onnxruntime-openvino · gllm-inference-binary · onnxruntime-genai · optimum · vllm · nncf