Packages
Provides a Python API for annotating events and code ranges in applications to capture performance profiling data visible in NVIDIA's Visual Profiler.
Install only if you have CUDA and Visual Profiler available; otherwise it provides no value.
Provides deprecated integration for creating and querying indexes on LlamaCloud's managed parsing and retrieval service, now superseded by the llama-cloud package.
Strands Agents is a Python SDK for building AI agents that work with multiple model providers and integrates with Model Context Protocol (MCP) servers to access pre-built tools.
Python SDK that traces, evaluates, and analyzes LLM application behavior through automatic and manual instrumentation, with built-in support for OpenTelemetry and popular LLM libraries.
Install it if you need tracing and evaluation for LLM products and have access to a Laminar backend; skip it if you have no observability backend or prefer a…
ml_dtypes provides NumPy-compatible data types for machine learning, including bfloat16, multiple 8-bit and 4-bit floating-point formats, and narrow integer encodings (1, 2, and 4-bit).
Connects LangChain applications to Google Cloud's Vertex AI generative models, enabling use of Google's language models, embeddings, and vector search within LangChain workflows.
Install it if you're building LangChain applications on Google Cloud and want native Vertex AI model access.
Computes embeddings and reranking scores for text using pre-trained transformer models, enabling semantic search, similarity matching, and ranking tasks.
Install it if you need embeddings or reranking; the only gotcha is ensuring your environment has torch and the required Python version (3.10+).
Python SDK for building and interacting with Claude Code agents, providing async query functions, custom tool definition, and hook-based control over agent behavior.
Einops provides readable tensor operations (rearrange, reduce, repeat, pack, unpack, einsum) that work across numpy, PyTorch, TensorFlow, JAX, and other frameworks without framework-specific syntax.
Install it if you work with tensors across any supported framework.
LightGBM is a gradient boosting framework for classification, regression, and ranking tasks that uses tree-based learning algorithms optimized for speed and memory efficiency.
SWE-bench is a benchmark framework for evaluating language models on real-world GitHub issues by having them generate patches that resolve described problems in actual codebases.
Install only if you have or can access Docker and sufficient compute resources; the package itself installs easily, but meaningful evaluation requires significant…
Accelerate abstracts away distributed training boilerplate for PyTorch, letting you run the same training script on single CPU, single GPU, multi-GPU, TPU, or mixed-precision setups with minimal code changes.
Provides NVIDIA CUBLAS native runtime libraries for CUDA 12, enabling GPU-accelerated linear algebra operations in Python applications.
Provides NVIDIA CUDA NVRTC native runtime libraries for compiling CUDA code at runtime on Windows and Linux systems.
Provides NVIDIA CUSPARSE native runtime libraries for GPU-accelerated sparse matrix operations on CUDA 12 systems.
Provides cuDNN runtime libraries for GPU-accelerated deep neural network operations, requiring an NVIDIA GPU and CUDA 12 environment.
Provides NVIDIA's JIT LTO compiler library for CUDA 12, enabling just-in-time compilation and link-time optimization in GPU-accelerated applications.
Provides NVIDIA CUFFT native runtime libraries for CUDA 12, enabling GPU-accelerated Fast Fourier Transform computations in Python applications.
However, verify NVIDIA's proprietary license terms for your use case, and ensure your system has compatible NVIDIA hardware and drivers installed before proceeding.
wandb is a machine learning experiment tracking and visualization platform that logs metrics, hyperparameters, datasets, and models to a centralized dashboard for monitoring and comparing training runs.
Install it if you need centralized experiment tracking, comparison, and visualization for machine learning projects; skip it if you prefer local-only logging or have…
Provides CUDA solver native runtime libraries for GPU-accelerated linear algebra operations on NVIDIA hardware.
Provides GPU profiling runtime libraries for CUDA 12, enabling third-party tools to access NVIDIA GPU profiling APIs.
Install only if a higher-level tool or framework explicitly lists it as a dependency.
Provides NVIDIA CURAND native runtime libraries for CUDA 12, enabling GPU-accelerated random number generation on supported hardware.
However, verify that your application actually requires it (it is typically pulled in transitively by higher-level libraries) and that you have an NVIDIA GPU and CUDA…
Provides NVIDIA CUDA 12 runtime native libraries for GPU-accelerated computing on Linux, Windows, and ARM64 systems.
Provides a C-based API for annotating events, code ranges, and resources in applications to enable capture and visualization through NVIDIA's Visual Profiler.
However, the aging maintenance status (435 days since last release) suggests limited active development—verify that CUDA 12.x is your target version and that the…
Provides LangChain integration for Anthropic's Claude models, enabling developers to use Claude through the LangChain framework for generative AI applications.
Install it if you're already using LangChain and want to use Claude; it's the standard way to do so within the LangChain ecosystem.
JAX is a Python library for automatic differentiation, XLA compilation, and function transformation of numerical code, enabling high-performance machine learning and scientific computing on CPUs, GPUs, and TPUs.
Install it if your workflow involves numerical computing, machine learning research, or scientific computing on accelerators; avoid it only if you need only basic…
Instrument GenAI applications with distributed tracing to log execution flows, spans, and metadata to MLflow backends for observability and monitoring.
Install only if you have access to a remote MLflow backend (Databricks, SageMaker, Nebius, or self-hosted); local-only tracing is not supported.
Provides OpenTelemetry semantic conventions and attributes for instrumenting generative AI applications, enabling structured logging of prompts, completions, token usage, and other LLM-specific telemetry.
Provides legacy LangChain chains, community integrations, indexing APIs, and deprecated functionality for backward compatibility with older LangChain applications.
SageMaker Python SDK is a library for training and deploying machine learning models on Amazon SageMaker, supporting frameworks like PyTorch and MXNet as well as Amazon's built-in algorithms.
TensorFlow is an open-source machine learning framework for building and training neural networks and other numerical computation models, deployable across CPUs, GPUs, TPUs, and edge devices.
Provides legacy registered functions and model architectures for backwards compatibility with older configuration files.
However, its abandoned status since 2023-01-23 means no bug fixes or security patches will be provided.
ONNX provides an open-source format and runtime for representing and executing AI models across different frameworks and hardware platforms, enabling model interoperability and inference.
Install it if you need to work with ONNX models, export models to ONNX format, or build cross-framework inference pipelines.
Provides a Python client to interact with Ollama, a local or cloud-based large language model server, enabling chat, text generation, embeddings, and model management through a simple API.
Install it if you are already running Ollama or plan to; if you have no Ollama server, it is not useful on its own.
Instructor wraps LLM APIs to extract validated, typed structured data from natural language by defining Pydantic models and letting the package handle schema generation, validation, and retries.
Install it if you need to extract validated structured data from LLM responses; skip it if you're building agents or need richer observability (the docs suggest…
SHAP computes Shapley values to explain individual predictions and feature importance across any machine learning model, using game-theoretic attribution to show how each feature contributes to model output.
Gradio builds web interfaces for machine learning models, APIs, and Python functions without requiring JavaScript or web hosting knowledge, then shares them via public URLs.
Docling parses diverse document formats—PDF, DOCX, PPTX, XLSX, HTML, EPUB, email, images, video, audio, and more—into a unified representation, with advanced PDF layout understanding and export to Markdown, HTML, JSON, and other formats.
Integrates Google's Gemini AI models (chat, vision, embeddings) into LangChain applications through a unified interface.
Optuna is a hyperparameter optimization framework that automates the search for optimal hyperparameter values in machine learning models using a define-by-run API and state-of-the-art sampling algorithms.