Packages
PaddleX is a low-code framework for training, inference, and deployment of computer vision and document processing models, integrating pre-trained models across OCR, object detection, image classification, segmentation, and time-series tasks.
Performs optical character recognition (OCR) on images to extract text, supporting Chinese, English, and other languages via ONNX Runtime inference.
Install it if you need offline OCR with reasonable speed and accuracy; verify model download behavior and performance on your target hardware before committing to…
Provides cuDNN runtime libraries for GPU-accelerated deep neural network operations on NVIDIA CUDA 11 hardware.
Decord decodes video and audio files with hardware-accelerated decoders (FFMPEG, Nvidia, Intel codecs) and provides efficient random frame access for deep learning workflows.
Install only if your video formats and Python version are confirmed compatible, and consider alternatives if active maintenance is critical for your project.
PaddlePaddle is a deep learning framework for building, training, and deploying neural networks across multiple platforms and devices.
Python SDK for transcribing and understanding audio using AssemblyAI's AI models, supporting prerecorded files, URLs, real-time streaming, and synchronous batch transcription.
A generated Python client library for the Neptune API, providing synchronous and asynchronous methods to call Neptune endpoints with automatic model parsing and optional authentication.
Provides helper functions to process images and videos for use with Qwen-VL vision-language models, handling multiple input formats including local files, URLs, base64-encoded data, and PIL images.
Provides an MCP server that exposes the Lean theorem prover to LLM agents via the Language Server Protocol, enabling programmatic access to proof states, diagnostics, and theorem search tools.
Install it if you use Cursor, Claude Code, or similar platforms with Lean projects.
Queries and fetches experiment metadata, runs, and attributes from Neptune ML tracking projects as pandas DataFrames, with filtering support.
LlamaIndex Legacy is a data framework for building LLM applications that connects private data to language models through indexing, retrieval, and query interfaces.
However, verify whether migration to the current package is recommended for your use case, as this is explicitly a legacy release.
Connects Groq's inference API to LangChain, enabling language model applications to use Groq as a provider for chat and text generation tasks.
Provides a LangChain callback handler to log LangChain executions to Braintrust for tracing and evaluation—now deprecated in favor of the integration built into the main braintrust package.
Chex provides utilities for writing reliable JAX code, including assertions for tensor properties, debugging helpers, and test variants to validate code across different JAX execution modes.
Install it if you write JAX code and want better visibility into shape, dtype, and execution-mode issues.
Provides the NVVM compiler IR library for building and optimizing CUDA applications on NVIDIA GPUs.
Speaker diarization toolkit that identifies and separates individual speakers in audio files using PyTorch-based pretrained models, with options for local open-source or cloud-hosted premium pipelines.
The main gotcha is the ffmpeg system dependency and the 21-package dependency tree—typical for PyTorch-based audio ML but not lightweight.
Provides Amazon Bedrock embedding models integration for LlamaIndex, enabling text-to-vector conversion using Titan and Cohere embeddings through AWS.
Provides CUDA 11 runtime native libraries for GPU-accelerated computing on Linux and Windows systems.
Unsloth accelerates training and fine-tuning of large language models and diffusion models on consumer hardware, reducing memory usage and training time through optimized PyTorch integration.
Outlines constrains LLM generation to produce structured outputs matching a specified type—JSON schemas, Pydantic models, enums, or literals—without post-hoc parsing.
Install it if you need reliable structured outputs from language models without post-hoc parsing.
Agno is a framework and runtime for building, deploying, and managing agent platforms with built-in APIs, storage, security, and observability.
Client library for posting experiment metrics, model registry data, and authentication to DVC Studio, a platform for tracking and managing machine learning experiments.
Provides NVIDIA CUDA NVRTC (runtime compilation) native libraries for Python, enabling just-in-time compilation of CUDA kernels on x86_64 and aarch64 Linux and Windows systems.
Tecton is a Python SDK for defining, testing, and deploying machine learning feature pipelines that compute and serve features for training and real-time inference.
Provides dict-like configuration data structures with dot-notation access, type safety, and lazy evaluation for ML experiment and model configuration.
XFormers provides reusable, composable building blocks for constructing Transformer neural network architectures, enabling flexible assembly of state-of-the-art model variants without monolithic implementations.
Provides a shared Python API layer for building agents and applications that integrate with Databricks AI features like Vector Search and AI/BI Genie.
Downloads and loads pre-trained TensorFlow SavedModels from TensorFlow Hub (now redirected to Kaggle Models) for reuse in TensorFlow programs with minimal code.
However, be aware that maintenance is dormant and many models from the original tfhub.dev have been deleted; verify your models exist before committing to this…
Defines and optimizes tunable hyperparameter pipelines using Optuna-backed search, allowing you to compose processing workflows with searchable parameters and automatic tuning against a loss function.
However, the aging maintenance status (339 days since last release) and unclear license are cautions—verify the license before production use and expect slower issue…
Cog packages machine learning models into production-ready Docker containers with automatic CUDA/dependency resolution, OpenAPI schema generation, and a built-in HTTP inference server.
Grain is a Python library for reading, transforming, and batching data for training and evaluating machine learning models, with support for declarative data processing pipelines.
Provides PyTorch-native audio data augmentation transforms that run on CPU or GPU, designed to integrate directly into neural network models as differentiable modules.
However, be aware that maintenance is aging (last release 576 days ago), multiprocessing and multi-GPU setups have known limitations, and some transforms have edge…
fvcore provides shared computer vision utilities for PyTorch, including neural network layers, loss functions, FLOP counting, parameter analysis, and hyperparameter scheduling used across FAIR's research frameworks.
However, note that the last PyPI release was December 2022—verify compatibility with your PyTorch version before installing, and consider installing from GitHub if…
Provides PyTorch-based filterbank implementations for audio signal processing, particularly designed for audio source separation tasks.
However, be aware that the project is dormant—no active maintenance or updates are expected, so compatibility with very recent PyTorch or Python versions is not…
Loads safetensors model files significantly faster than the standard safetensors deserializer by optimizing I/O patterns for GPU and storage systems.
Provides NVIDIA CUDA C Runtime libraries for Python, enabling GPU-accelerated compute on supported platforms.
Install only if you have a specific dependency requirement from another package or a documented need for CUDA 13.3 runtime support; it is not typically installed…
Adds the What-If Tool, an interactive visual interface for exploring and understanding ML model behavior, as a plugin to TensorBoard.
However, do not rely on this package for new features or bug fixes; it is abandoned and has not been updated since 2022-01-05.
Optimum provides optimization tools to export and run Transformers, Diffusers, and other HuggingFace models efficiently on specialized hardware accelerators like ONNX Runtime, OpenVINO, AWS Trainium, and Intel Gaudi.
Generates visual diagrams of PyTorch neural network computation graphs and autograd traces, showing layer connections and tensor flow through the model.
However, maintenance is dormant (620 days since last release), so verify compatibility with your PyTorch version before relying on it for production workflows.
GPU-accelerated attention mechanism implementation using CuTeDSL for Hopper and Blackwell GPUs, optimizing the core computational bottleneck in transformer models.
However, it is in alpha status and only useful for specific GPU hardware; it will not benefit users on other GPU architectures or CPU-only setups.