Packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
LangChain Core provides the foundational abstractions and interfaces for building LLM applications, enabling modular composition of language model chains, agents, and tools across the LangChain ecosystem.
Install it if you are building any LLM application that benefits from modular abstractions and ecosystem interoperability.
Loads and preprocesses datasets from the Hugging Face Hub or local files in many formats (CSV, JSON, Parquet, Arrow, audio, image, video, PDF, NIfTI), with built-in support for streaming, caching, and conversion to NumPy, Pandas, PyTorch, TensorFlow, and other frameworks.
Install it if you work with datasets for machine learning, data exploration, or preprocessing—it will save time and reduce boilerplate.
LangSmith is a Python client SDK for the LangSmith observability and evaluation platform, enabling you to trace, log, and evaluate language model applications and agents.
Install it if you build LLM applications and need centralized tracing, debugging, and evaluation—especially if you already use LangChain.
Serializes and deserializes tensors to and from a safe, standardized binary format designed for secure model storage and sharing.
Install it if you work with tensor serialization or model checkpoints.
PyTorch provides GPU-accelerated tensor computation and automatic differentiation for building and training deep neural networks in Python.
SGLang is a serving framework that runs large language models and multimodal models on GPUs and TPUs with optimizations for low-latency, high-throughput inference across single and distributed setups.
Install it if your workload involves serving models at scale; skip it if you only need local, single-model inference without distributed optimization.
SDK client for LlamaCloud services: document parsing (LlamaParse), structured data extraction (LlamaExtract), and automated document ingestion with retrieval (LlamaCloud Index).
onnxruntime loads and executes Open Neural Network Exchange (ONNX) models with a focus on inference performance across CPUs and accelerators.
Install it if you have ONNX models to run in production or development.
FastMCP is a Python framework for building Model Context Protocol (MCP) servers, clients, and interactive applications that connect large language models to tools, resources, and data.
FastMCP exposes Python functions as Model Context Protocol (MCP) tools, resources, and prompts that LLMs can call, with automatic schema generation, validation, and documentation.
Install it if you need to expose Python functions to LLMs or integrate with MCP-compatible tools.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Provides LangChain integrations for OpenAI's API, enabling language models, embeddings, and other OpenAI services to be used within LangChain applications.
Install it if you're building a LangChain application and want to use OpenAI as your model provider—it's the standard way to do so.
Provides Python type definitions (TypedDict, Literal) for the LangChain agent streaming protocol wire format, enabling type-safe integration with protocol-compliant clients and servers.
XGBoost is a gradient boosting library that trains tree-based machine learning models for classification, regression, and ranking tasks, with support for distributed computing across Kubernetes, Hadoop, Spark, and Dask.
Provides NVIDIA's NCCL runtime library for GPU collective communication operations including all-reduce, all-gather, reduce, broadcast, and reduce-scatter.
Install only if your system has compatible NVIDIA GPUs and CUDA 12 already installed.
Python client SDK for the Mistral AI API, providing access to chat completions, embeddings, file uploads, and agent interactions via synchronous and asynchronous interfaces.
Splits text documents into chunks using a variety of strategies, designed to work with LangChain's language model pipelines.
Torchvision provides pre-built datasets, model architectures, and image transformation utilities for computer vision tasks within the PyTorch ecosystem.
mlflow-skinny is a lightweight client library for MLflow that enables experiment tracking, model logging, and observability for AI and ML applications without bundling SQL storage, a server, or data science dependencies.
Provides NVIDIA CUBLAS native runtime libraries for GPU-accelerated linear algebra operations in CUDA-enabled environments.
Provides NVIDIA CUDA NVRTC (NVIDIA Runtime Compilation) native runtime libraries for Python, enabling runtime compilation of CUDA kernels on x86_64 Linux, ARM64 Linux, and Windows platforms.
Provides NVIDIA's JIT LTO compiler library for Python, enabling just-in-time and link-time optimization compilation functionality for CUDA-based applications.
Provides cuDNN runtime libraries for GPU-accelerated deep neural network primitives, requiring CUDA 13 and nvidia-cublas as dependencies.
Install only if you have CUDA 13 and nvidia-cublas available; otherwise, installation will not resolve the underlying GPU dependencies.
Provides NVIDIA's Collective Communication Library (NCCL) runtime for GPU-accelerated collective operations like all-reduce, all-gather, reduce, broadcast, and reduce-scatter across multiple GPUs using PCIe, NVLink, NVswitch, InfiniBand, or TCP/IP.
However, verify that your framework (PyTorch, TensorFlow, etc.) declares it as a dependency rather than installing it standalone.
MLflow is an open-source platform for managing the complete lifecycle of machine learning and AI applications, including experiment tracking, model evaluation, deployment, LLM observability, and prompt management.
NVSHMEM provides a global address space for GPU cluster communication, enabling fine-grained GPU-initiated and CPU-initiated operations across multiple GPUs' memory via a parallel programming interface based on OpenSHMEM.
However, the unclear license status and lack of public documentation are concerns—verify licensing terms and review NVIDIA's CUDA Zone documentation before committing…
Provides NVIDIA CUSPARSE native runtime libraries for GPU-accelerated sparse matrix operations in CUDA applications.
Install only if you have NVIDIA GPU hardware and the CUDA toolkit already set up; this is a runtime library, not a standalone tool.
Provides NVIDIA CUFFT native runtime libraries for GPU-accelerated Fast Fourier Transform computations on CUDA-enabled hardware.
Provides CUDA solver native runtime libraries for GPU-accelerated linear algebra operations on NVIDIA hardware.
Provides NVIDIA CURAND native runtime libraries for GPU-accelerated random number generation in Python on Linux, Windows, and aarch64 platforms.
Provides NVIDIA CUDA Runtime native libraries for GPU-accelerated computing on Windows and Linux systems.
Provides CUDA profiling runtime libraries that enable third-party tools to access GPU profiling APIs on NVIDIA hardware.
However, it requires an NVIDIA GPU and CUDA environment, and the unclear license should be confirmed before commercial deployment.
Integrates OpenAI's language models into applications using LlamaIndex, providing access to completion, chat, and streaming APIs through a unified interface.
Install it if you are building with llama-index-core and want to use OpenAI models; it is the standard integration point for that use case.
Provides Python bindings for NVIDIA's cuFile GPUDirect storage access libraries, enabling direct GPU-to-storage I/O without CPU involvement.
However, verify the unclear license terms before use in production, and confirm that your system has the required cuFile runtime libraries installed.