Packages
LlamaIndex is a data framework that connects large language models to your own data sources, enabling retrieval-augmented generation (RAG) and agentic applications through data connectors, indexing, and query interfaces.
Provides a unified format for storing and loading compressed neural network tensors, supporting multiple quantization and sparsity schemes like GPTQ, AWQ, SmoothQuant, INT8, and FP8.
Renders and parses the Harmony response format used by OpenAI's gpt-oss models, enabling structured conversation handling, reasoning output, and function calls in Python.
Provides type annotations and runtime type-checking for array shape and dtype across JAX, PyTorch, NumPy, MLX, and TensorFlow, with no JAX dependency required.
Reads and writes binary files in the GGUF (GGML Universal File) format, commonly used for storing machine learning models compatible with GGML-based tools like llama.cpp.
Install it if you need to work with GGUF model files programmatically or via command-line tools.
GEPA optimizes textual system components—prompts, code, agent architectures, configurations—using LLM-based reflection and Pareto-efficient evolutionary search to improve performance against any measurable metric.
Provides supplementary files for the LiteLLM Proxy AI Gateway, reducing the size of the main litellm package by separating optional proxy server dependencies.
Deep Agents is an opinionated agent harness that runs LLM-powered agents out of the box, with built-in filesystem access, context management, sub-agent delegation, and skill loading—extensible at any layer without forking.
Provides Google Cloud Storage filesystem support for TensorFlow, enabling direct reading and writing of data from GCS buckets within TensorFlow pipelines without local downloads.
Install only if your TensorFlow version matches the compatibility table (0.37.1 requires TensorFlow 2.16.x).
Profiles PyTorch models by counting Multiply-Accumulate Operations (MACs) and parameters in a single forward pass, with built-in rules for common layer types and support for custom counting rules.
Install it if you need to profile PyTorch model efficiency.
DSPy is a framework for building and optimizing modular language model systems through Python code rather than manual prompt engineering, with built-in algorithms for teaching models to produce high-quality outputs.
LiteLLM Enterprise provides a unified Python SDK and self-hosted AI Gateway to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) through a single OpenAI-compatible interface.
However, the proprietary license is not clearly documented in the package metadata—verify licensing terms before deploying in production or in open-source projects.
Lightning is a framework that organizes PyTorch code to automate distributed training infrastructure—handling backpropagation, mixed precision, multi-GPU, and multi-node setups—while keeping your model logic unchanged.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
CatBoost is a gradient boosting library that trains decision-tree ensembles for classification, regression, and ranking, with built-in support for categorical features and GPU acceleration.
Install it if you work with tabular data and want gradient boosting with native categorical feature support, GPU acceleration, or distributed training.
Provides ready-to-use tools for AI agents to perform file operations, web search, AWS integration, Slack messaging, Python code execution, and browser automation.
Roboflow is a Python client for the Roboflow computer vision platform, enabling you to create projects, upload datasets and images, train vision models, and run inference on models hosted on Roboflow or self-hosted via Roboflow Inference.
Install it if you are already using Roboflow's platform and want to automate project and model management from Python.
vLLM is a high-throughput inference and serving engine for large language models, offering optimized memory management, continuous batching, and support for multiple hardware platforms and model architectures.
Install it if you're building an LLM application, API service, or batch inference pipeline; skip it if you only need simple single-model inference without serving…
Unified client library for chat completions, text embeddings, and image embeddings across GitHub Models, Azure AI Foundry deployments, and Azure OpenAI Service.
However, it is still in beta (1.0.0b9), so expect potential API changes before a stable release.
Detects sentence boundaries in text using rule-based heuristics, splitting paragraphs into individual sentences while handling edge cases like abbreviations and decimal points.
However, do not install if you need active maintenance, multi-language support, or assurance of ongoing bug fixes.
TensorFlow Text provides text preprocessing operations and tokenizers that run within the TensorFlow computation graph, enabling consistent text handling across training and inference without external preprocessing scripts.
Install it if you are building NLP models with TensorFlow and need tokenization or text normalization; ensure your tensorflow-text version matches your TensorFlow…
Orbax Checkpoint provides asynchronous checkpointing for JAX machine learning workflows, supporting multiple storage formats and customizable serialization to save and restore model state during training.
Install it if you're running JAX training jobs that need reliable state persistence.
Docling Slim is a lightweight, modular SDK for parsing and converting documents (PDF, DOCX, HTML, Markdown, and others) into a unified representation, with optional extras for specific formats and features.
Provides Python APIs for loading, parsing, and working with the MS-COCO dataset, including annotation access and evaluation metrics for object detection and segmentation tasks.
Enforces structured output from large language models by computing token masks that constrain decoding to valid context-free grammars, JSON schemas, or regular expressions with minimal per-token overhead.
Ingests and pre-processes unstructured documents (PDFs, HTML, Word, emails, images) into structured elements for machine learning pipelines, supporting 60+ file types with parsing, chunking, and enrichment.
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
Provides Python client APIs to communicate with TensorFlow Serving, a production machine learning model serving system using gRPC for deployment and inference.
Install only if you already have or plan to run a TensorFlow Serving instance; it is a client library, not a standalone serving system.
ModelScope provides unified Python interfaces for inference, fine-tuning, and evaluation across machine learning models spanning NLP, computer vision, speech, multi-modal, and scientific computing domains.
Outlines-core provides structured text generation for language models by converting JSON schemas into finite-state automata that constrain token selection during generation.
FlashInfer provides optimized GPU kernels for LLM inference operations—attention, GEMM, and mixture-of-experts—with support for multiple GPU architectures and low-precision quantization.
PyOD detects anomalies and outliers across tabular, time series, graph, text, image, and audio data using 61 detectors, with optional agentic workflows for AI-driven investigation.
Generate access tokens and call LiveKit server APIs (room management, egress, ingress, SIP, agent dispatch, connectors) from Python backends using async/await.
Install it if you are building a backend that needs to generate tokens, manage rooms, or integrate with LiveKit's services.
SWE-smith generates large-scale software engineering training datasets by synthesizing task instances from GitHub repositories and managing their execution environments via Docker.
Provides the pipeline specification and protobuf definitions for Kubeflow Pipelines, enabling serialization and validation of ML workflow configurations on Kubernetes.
Official Python SDK for xAI's APIs, providing synchronous and asynchronous clients to interact with text generation, image understanding, video generation, and structured output models.
Provides an ahead-of-time compiled module for faster DLPack v1.2 conversion with torch, avoiding JIT compilation overhead and compiler toolchain requirements.
However, maintenance is aging (last release 214 days ago), so verify compatibility with your specific torch and library versions before relying on it in production.
Python SDK for building, deploying, and managing AI agents and generative AI applications on Google's Gemini Enterprise Agent Platform (formerly Vertex AI).
JiWER computes speech recognition evaluation metrics (WER, MER, WIL, WIP, CER) by calculating minimum-edit distance between reference and hypothesis text using RapidFuzz for speed.
Autoevals provides automatic evaluation methods for AI model outputs, including LLM-as-a-judge, heuristic, and statistical approaches, with built-in support for subjective tasks like fact-checking and safety assessment.