Packages
Provides Python bindings to Pinecone's Assistant APIs for creating, managing, and chatting with AI assistants that can ingest and reason over documents.
Install it if you're already using Pinecone and want to leverage its managed Assistant APIs.
Provides optimized MLA (Multi-head Latent Attention) kernels for NVIDIA Blackwell GPUs, including prefill and decode operations with FP8 quantization and ragged sequence support.
OpenVINO converts and optimizes deep learning models from various frameworks for inference on CPUs, GPUs, and AI accelerators without requiring the original training frameworks.
Integrates ElevenLabs text-to-speech into the LiveKit Agents framework for building realtime voice agents that can speak with ElevenLabs' voices.
Composio provides a Python SDK that connects AI agents to 1000+ pre-authenticated third-party app toolkits, managing per-user sessions, OAuth flows, event triggers, and sandboxed execution.
Provides common utilities and functions for converting machine learning models from various AI frameworks to ONNX format, enabling interoperability between different framework converters.
Install it if you are building or using ONNX converters, especially when working with multiple frameworks or planning to leverage existing converter ecosystems.
Extracts text, structure, and key-value pairs from documents using Azure's cloud-based machine learning models, supporting both prebuilt document types and custom trained models.
Install only if you are committed to using Azure's Document Intelligence service; it is not a standalone solution and requires cloud resources and API credentials.
Provides NVIDIA CUDA C++ Core Compute Libraries (CCCL) for GPU-accelerated computing on Windows and Linux systems.
However, verify the unclear license terms with NVIDIA and confirm your system architecture matches an available wheel before installing.
A Python SDK for accessing 400+ AI models across multiple providers through the OpenRouter API, with type-safe request handling and support for both synchronous and asynchronous operations.
TensorFlow CPU provides a machine learning and numerical computation framework for building and training models on CPU-based systems, with support for Python 3.10 through 3.13.
A PyTorch VRAM allocator that dynamically offloads model weights to system memory when GPU memory is under pressure, using virtual address reservation to minimize fragmentation.
A Python client for running machine learning models on Replicate's API, supporting synchronous and asynchronous execution, file streaming, and webhook-based background jobs.
Install it if you need to call Replicate models from Python—it's the standard integration point.
Implements vector quantization layers for PyTorch, enabling discrete codebook-based compression of continuous embeddings used in generative models.
Solves stochastic differential equations (SDEs) with GPU support and efficient backpropagation through PyTorch, enabling gradient-based learning of SDE-based models.
CMA-ES is a derivative-free numerical optimization algorithm for difficult non-convex, multi-modal, or ill-conditioned optimization problems in continuous or mixed-integer search spaces.
Install it if you need derivative-free optimization for difficult non-convex problems in moderate dimensions; skip it if you have gradients available or are…
OpenHands is a self-hosted developer control center that runs coding agents (OpenHands, Claude Code, Codex, Gemini, or any ACP-compatible agent) across local, remote, and cloud backends, with built-in support for automations and integrations.
Connects Hugging Face models and embeddings to LangChain applications, providing integration classes for local inference and API-based access to Hugging Face resources.
DBOS adds durable workflows and queues to Python applications by checkpointing execution state in Postgres, allowing programs to automatically recover from failures without external orchestration infrastructure.
Provides access to many public datasets as tf.data.Datasets, handling download, preparation, and loading for TensorFlow workflows.
Install it if you work with TensorFlow and need quick access to public datasets; skip it only if you manage datasets entirely through custom pipelines.
SpeechBrain is a PyTorch-based toolkit for building speech and text processing systems, providing pretrained models, training recipes, and inference interfaces for tasks like speech recognition, speaker identification, speech enhancement, and language modeling.
Milvus Lite is a pure-Python local vector database that provides dense and sparse vector search, BM25 full-text search, and scalar filtering through a Milvus-compatible API, storing data in a local `.db` file or embedded gRPC server.
Exposes Unity Catalog functions as LangChain tools, allowing agents and LLMs to call UC functions directly within agentic workflows.
Solves the linear assignment problem using the Jonker-Volgenant (LAPJV) or Volgenant-Mordecai (LAPMOD) algorithm, returning optimal row-to-column assignments for dense or sparse cost matrices.
Install it if you need to solve linear assignment problems and prefer a specialized implementation over a general-purpose optimizer.
Python client library for the AIStudio API, enabling developers to build projects, applications, and plugins using AIStudio models and the Yiyan API.
Provides Python APIs to create, retrieve, and execute Unity Catalog functions with both async and sync interfaces, supporting integration of UC functions as tools in GenAI agents.
Install it if you need to integrate Unity Catalog functions into Python applications or GenAI agent workflows.
Graphiti builds and queries temporal context graphs for AI agents, tracking how facts change over time with full provenance to source data, supporting both semantic and keyword retrieval alongside graph traversal.
Provides core APIs and utilities for managing Azure Machine Learning workspaces, experiments, compute resources, datasets, models, and training runs.
However, do not start new projects with it—it is deprecated and will receive only security fixes until June 2026.
Ragas provides objective metrics, test data generation, and evaluation workflows for Large Language Model applications, integrating with frameworks like LangChain and OpenAI.
However, verify the two known vulnerabilities (GHSA-95ww-475f-pr4f, PYSEC-2026-3046) for your threat model, and be aware that the 19 runtime dependencies will add…
Converts ONNX model files to LiteRT, TensorFlow, PyTorch, TorchScript, and other formats, with support for direct conversion from LiteRT back to PyTorch.
Connects LangChain applications to Chroma, a vector database, enabling semantic search and retrieval-augmented generation workflows.
Install it if you are already using LangChain and want to use Chroma as your vector store; it is the standard way to connect the two.
Unified framework for evaluating generative language models against over 60 standard academic benchmarks with hundreds of task variants, supporting multiple model backends and inference engines.
Install only if you have a model backend in mind and Python >=3.10.
Fiddle is a Python-first configuration library that lets you define and manage complex program parameters in readable Python code, particularly suited for machine learning applications.
Install it if configuration-as-code in Python appeals to your workflow; skip it if you prefer external config files or simpler parameter passing.
BrowserGym provides a gymnasium environment for automating web tasks in Chromium, letting you build and test agents that interact with websites through a standardized interface.
Not recommended for production web scraping or automation—it is explicitly a research tool.
Provides NVIDIA CUFFT native runtime libraries for CUDA 11, enabling GPU-accelerated Fast Fourier Transform computations on compatible systems.
Install only if required as a transitive dependency of an active package, or if you are locked to CUDA 11 and cannot upgrade.
Provides CUSPARSE native runtime libraries for CUDA 11, enabling sparse matrix operations on NVIDIA GPUs.
Provides NVIDIA CUDA solver native runtime libraries for GPU-accelerated linear algebra and matrix operations on CUDA 11 hardware.
nemo-toolkit provides a PyTorch framework for building, training, and deploying speech AI models including automatic speech recognition (ASR), text-to-speech (TTS), and speech-based language models.
Provides NVIDIA CURAND native runtime libraries for CUDA 11, enabling GPU-accelerated random number generation in Python applications on x86_64 and ARM64 Linux, Windows, and legacy x86 platforms.
No—install only if you are locked into CUDA 11 and have no path to upgrade.
Provides NVIDIA CUDA profiling runtime libraries (CUPTI) for GPU performance analysis and monitoring on systems with CUDA 11 support.
Builds and orchestrates AI agents powered by OpenAI and Azure OpenAI APIs, with support for multi-agent workflows, tool calling, and agent-to-agent collaboration patterns.