Packages
Flax is a neural network library for JAX that lets you build, train, and debug deep learning models using Python objects with reference semantics instead of functional transformations.
Not recommended if you prefer a more batteries-included framework or are new to JAX.
Provides Python and C++ APIs to NVIDIA's cuDNN library, exposing high-performance GPU kernels for scaled dot-product attention, grouped matrix multiplication for mixture-of-experts training, and fused operations optimized for Hopper and Blackwell GPUs.
Install it if you are training or deploying deep learning models on NVIDIA Hopper or Blackwell GPUs and need high-performance attention, grouped GEMM, or quantized…
A framework for building realtime multimodal and voice AI agents that connect to LiveKit rooms and handle audio/video interactions with language models.
Python SDK for building real-time video, audio, and data applications by connecting to LiveKit servers as a participant or managing rooms via server APIs.
Whisper performs multilingual speech recognition, speech translation, and language identification using a Transformer model trained on diverse audio data.
Install only if you need multilingual speech recognition or translation; for English-only transcription, smaller or specialized models may be more practical.
TorchCodec decodes and encodes videos, audio, and images to and from tensors on CPU and CUDA GPUs, wrapping FFmpeg for video/audio and providing native image codecs.
Provides a Python client library for interacting with Azure Machine Learning services, enabling job submission, pipeline orchestration, model management, and AutoML training across multiple task types.
Install it if you need to programmatically interact with Azure Machine Learning services—it is the official and recommended SDK for that purpose.
SacreBLEU computes BLEU, chrF, and TER scores for machine translation evaluation with automatic test set management and reproducible, comparable results across systems.
Install it if you work with machine translation evaluation, benchmark scoring, or need comparable BLEU/chrF/TER metrics across systems.
Download, upload, and manage AI models, datasets, and other assets on ModelScope Hub through a unified Python SDK and CLI interface.
This package is a deprecated stub that redirects users to install imbalanced-learn instead; it serves no functional purpose.
Adds Machine Learning Model extension support to PySTAC, enabling description of ML model architecture, framework, inputs, outputs, training details, and runtime information within STAC catalogs.
Install it if you are already using PySTAC and need to catalog machine learning models with standardized metadata.
Extends PySTAC to add support for the Label Extension specification, enabling description of labeled training data for machine learning with fields for label types, classes, tasks, methods, and label distribution overviews.
Converts TensorFlow models to TensorFlow.js format for browser and Node.js deployment, with CLI tools and a wizard for model conversion workflows.
Extracts medical entities and personally identifiable information from clinical text, then de-identifies it—all running locally on your hardware after required model artifacts are available.
semchunk splits text into semantically meaningful chunks while preserving local context, supporting custom tokenizers, chunk overlapping, offsets, and optional AI-powered chunking via the Isaacus API.
Install it if you need to chunk text for RAG, embeddings, or language model workflows.
TF-Keras is the pure-TensorFlow implementation of Keras, providing a high-level API for building and training deep learning models with TensorFlow as the backend.
Albucore provides optimized atomic image processing functions that automatically select the fastest implementation (NumPy, OpenCV, NumKong, StringZilla, or PyTorch) based on input characteristics for efficient uint8 and float32 image manipulation.
Install it if you are building image augmentation or preprocessing pipelines and want to avoid manual optimization.
Provides Python access to Voyage AI's embedding and reranking APIs for converting documents and queries into semantic vectors and scoring document relevance.
Install it if you need embeddings or reranking from Voyage AI; the API-only path has minimal overhead.
Mem0 adds a persistent, searchable memory layer to AI agents and assistants, storing and retrieving user preferences, session state, and interaction history across conversations.
EasyOCR performs optical character recognition on images across 80+ languages and writing scripts, extracting text with bounding boxes and confidence scores.
However, the last release was 689 days ago—verify compatibility with your PyTorch and Python versions before deploying to production, and monitor the repository for…
Detects faces and facial landmarks (eyes, nose, mouth) in images using a cascaded convolutional network, returning bounding boxes and keypoint coordinates.
It is suitable for production use in stable environments but not for projects requiring ongoing support.
SimSIMD provides SIMD-optimized kernels for computing vector distances, dot-products, and similarity measures across multiple data types and precisions, with support for spatial, probabilistic, and bit-level operations.
Install it if vector similarity or distance computation is a measurable bottleneck in your application.
Optax provides composable building blocks for gradient processing and optimization in JAX, including implementations of popular optimizers and loss functions that can be combined into custom solutions.
Install it if you are building machine learning systems with JAX and need flexible, composable optimizer and loss components.
Removes image backgrounds using deep learning models, available as a Python library, CLI tool, HTTP server, or Docker container.
RapidOCR extracts text from images using ONNX-based models optimized for speed and cross-platform deployment, supporting Chinese, English, and other languages with offline inference.
TRL provides trainer classes for post-training foundation models using techniques like supervised fine-tuning, direct preference optimization, and group relative policy optimization, built on transformers and accelerate.
Install it if you need to fine-tune, align, or distill language models with modern techniques; skip it only if you are working exclusively with inference or base…
TorchAO applies quantization and sparsity techniques to PyTorch models for faster training and inference with reduced memory usage, working natively with torch.compile() and FSDP2.
Opik is an open-source LLM observability and evaluation platform that logs traces of LLM calls, agents, and pipelines, then evaluates them with datasets, experiments, and LLM-as-a-judge metrics.
Provides metric learning loss functions, miners, and evaluation tools for training deep neural networks to learn embeddings where similar items cluster together and dissimilar items separate.
However, the last release was 362 days ago; if you need very recent bug fixes or features, verify that the current version addresses your use case or check the…
dlib-bin provides pre-compiled binary wheels of dlib, a machine learning and computer vision toolkit, eliminating the need to build from source during installation.
However, verify that the Boost Software License is compatible with your project, and confirm that the pre-built wheels meet your performance and feature requirements…
Provides Silero voice activity detection integration for the LiveKit Agents framework to detect when users are speaking in real-time voice agent applications.
Write ONNX functions and models in Python syntax, then convert them to ONNX graphs; includes tools for optimization and pattern-based graph rewriting.
The eager-mode debugger is a real productivity gain for model development, though not for production inference.
Integrates OpenAI's Realtime, Responses, LLM, TTS, and STT APIs into LiveKit Agents, plus support for OpenAI-compatible providers like Azure OpenAI, Cerebras, Fireworks, Perplexity, and others.
Install it if you're already using LiveKit Agents and want OpenAI integration.
Databricks Feature Engineering client for creating, managing, and serving feature tables within Databricks workspaces, including training models on feature data and publishing to online stores.
However, the proprietary license restricts use to Databricks Platform Services; confirm your Databricks agreement permits this library before production deployment.
Integrates Databricks AI services—LLMs, vector search, embeddings, and Genie—into LangChain applications through a unified package.
Install only if you have a Databricks workspace and LangChain is your chosen framework; it adds no value as a standalone tool.
Provides AI models for table structure recognition and page layout detection to support PDF document conversion.
Install friction is low, though torch and transformers are substantial dependencies—acceptable for ML workloads but not for lightweight applications.
Executes ONNX machine learning models on GPU hardware, providing inference acceleration for neural networks and other machine learning workloads.
Computes ROUGE scores (ROUGE-N, ROUGE-L, ROUGE-Lsum) to evaluate the quality of generated text summaries against reference summaries using n-gram overlap and longest common subsequence metrics.
Install it if you need ROUGE metrics for summarization evaluation and can tolerate high install friction from source-only distribution.
Python SDK for Alibaba Cloud Model Studio (Bailian) APIs, providing access to text generation, embeddings, image/video synthesis, speech processing, and multi-modal understanding models.
However, it is only useful if you have an Alibaba Cloud account and API key; it is not a general-purpose LLM client and does not support other model providers.
Integrates OpenAI language models with LlamaIndex agents, enabling agents to use OpenAI's models for reasoning and decision-making within the LlamaIndex framework.
However, the aging maintenance status (411 days since last release) warrants caution—verify compatibility with your target versions of llama-index-core and…