Packages
opencv-python provides pre-built CPU-only OpenCV bindings for Python, enabling computer vision tasks like image processing, object detection, and video analysis without requiring separate OpenCV installation.
OpenCV Python bindings for computer vision tasks, packaged without GUI dependencies for server and headless environments.
Install it if you need computer vision in a server or containerized environment where GUI dependencies are unnecessary.
tritonclient is a Python client library for communicating with Triton Inference Server over gRPC or HTTP, enabling you to send inference requests to remote model-serving deployments.
Install it if you are using Triton for model serving and need to query it from Python.
Torchmetrics provides a collection of PyTorch metrics implementations with automatic batch accumulation and multi-device synchronization, designed for distributed training workflows.
PyTorch Lightning wraps PyTorch training code to separate research logic from engineering boilerplate, enabling distributed training across GPUs, TPUs, and CPUs without code changes.
opencv-contrib-python provides Python bindings to OpenCV's full library, including contrib modules for advanced computer vision tasks like feature detection, image processing, and machine learning on images.
Ultralytics YOLO provides a unified framework for training and deploying computer vision models for object detection, instance segmentation, pose estimation, image classification, and semantic segmentation tasks.
Install it if you need a production-ready YOLO implementation and can accept the AGPL-3.0 license (or obtain a commercial license).
Profiles PyTorch models by counting Multiply-Accumulate Operations (MACs) and parameters in a single forward pass, with built-in rules for common layer types and support for custom counting rules.
Install it if you need to profile PyTorch model efficiency.
dlib-bin provides pre-compiled binary wheels of dlib, a machine learning and computer vision toolkit, eliminating the need to build from source during installation.
However, verify that the Boost Software License is compatible with your project, and confirm that the pre-built wheels meet your performance and feature requirements…
Provides OpenCV's computer vision algorithms and image processing functions in a headless (no GUI) package with contrib modules, designed for server and containerized environments.
Solves the linear assignment problem using the Jonker-Volgenant (LAPJV) or Volgenant-Mordecai (LAPMOD) algorithm, returning optimal row-to-column assignments for dense or sparse cost matrices.
Install it if you need to solve linear assignment problems and prefer a specialized implementation over a general-purpose optimizer.
nemo-toolkit provides a PyTorch framework for building, training, and deploying speech AI models including automatic speech recognition (ASR), text-to-speech (TTS), and speech-based language models.
Supervision provides utilities for loading, annotating, and processing computer vision datasets and model outputs—detection, segmentation, and classification results from any model framework.
Install it if you work with detection, segmentation, or classification models and want to avoid reinventing dataset utilities.
Augments images and related data (heatmaps, segmentation maps, keypoints, bounding boxes, polygons) by applying transformations like rotations, noise, cropping, and color shifts to expand training datasets for machine learning.
However, the package is dormant—last release was 2020-02-05, and two security vulnerabilities are recorded.
OCRmyPDF adds searchable text layers to scanned PDF files using Tesseract OCR, enabling them to be searched and copy-pasted while optionally deskewing, cleaning, and converting to PDF/A format.
Install it if you need to make scanned PDFs searchable.
Mlxtend provides ensemble methods, feature selection, visualization utilities, and frequent pattern mining algorithms for machine learning workflows.
Install it if you need these specific capabilities.
Reads, writes, and analyzes SVG Path objects and Bézier curves, providing geometric tools to transform, intersect, and measure path elements.
A Python client for the 2Captcha API that automates captcha solving across multiple captcha types including reCAPTCHA, FunCaptcha, GeeTest, Cloudflare Turnstile, and many others.
Automates machine learning model training and prediction on tabular data with minimal code, handling feature engineering, algorithm selection, and hyperparameter tuning internally.
tesserocr wraps Tesseract's C++ OCR engine via Cython, extracting text and metadata from images with support for Pillow objects and concurrent processing through Python's threading module.
AutoGluon Core provides the foundational infrastructure for automated machine learning, enabling training of high-accuracy models on tabular, time series, image, and text data with minimal code.
Automates feature engineering and preprocessing for machine learning pipelines, integrating with AutoGluon's broader ML automation framework to handle tabular, time series, and multimodal data preparation.
IceVision provides a unified framework for training and deploying object detection models, supporting multiple model architectures and training backends like PyTorch Lightning and Fastai.
Install only if you are maintaining legacy code already using IceVision and cannot migrate.
Python SDK for connecting to and running computer vision models and workflows on a local or remote Inference server, enabling image and video stream processing with object detection, classification, segmentation, and custom model inference.
Install it if you need to programmatically interact with Inference from Python; if you only need the server itself, install inference-cli instead.
ClearML Agent is a job scheduler and orchestration service that runs machine learning experiments on local or cloud resources, managing virtual environments, dependencies, and execution monitoring across Linux, macOS, and Windows.
Provides shared utilities and common infrastructure for AutoGluon's automated machine learning framework, supporting tabular, time series, multimodal, and image data tasks.
AutoGluon TimeSeries automates machine learning for time series forecasting, training and deploying high-accuracy models with minimal code using deep learning and statistical approaches.
A command-line tool for running computer vision inference locally via Docker or against Roboflow's hosted API, supporting object detection, classification, instance segmentation, and foundation models like CLIP and SAM.
However, verify that the GPL-3.0 and AGPL-3.0 licenses on bundled models (YOLOv5, YOLOv8) align with your project's licensing requirements before committing to…
AutoGluon automates machine learning model training and deployment across tabular, time series, image, text, and multimodal data with minimal code.
Install it if you need to train accurate ML models quickly across tabular, time series, or multimodal data without manual tuning.
FiftyOne is a Python framework for building, visualizing, and evaluating computer vision datasets and models, with integrated labeling, model evaluation, and data quality tools.
YOLOv5 is a packaged object detection model that runs inference on images and video to identify and localize objects, with integrated training, validation, export, and CLI support.
AutoGluon Multimodal automates machine learning on image, text, and mixed-data tasks, training and deploying high-accuracy models with minimal code using deep learning and foundation models.
Python wrapper for Intel RealSense SDK 2.0 that enables depth and color streaming from RealSense depth cameras, with access to calibration data and frame intrinsics.
Megatron Core provides GPU-optimized building blocks and parallelism strategies for training large transformer models at scale, including tensor parallelism, pipeline parallelism, and mixed precision support.
Not recommended for simple single-GPU training or inference-only use cases.
dlib is a C++ machine learning and computer vision toolkit with Python bindings, providing algorithms for image recognition, classification, and data analysis.
However, ensure your development environment has a working C++ compiler and CMake before attempting installation.
ETA is an extensible computer vision and machine learning analytics toolkit that provides core utilities for working with images, videos, embeddings, and ML inference pipelines, along with a CLI for building and running analytics workflows.
FiftyOne Brain provides AI/ML capabilities for analyzing and manipulating datasets and models, including visual similarity search, text-based querying, sample uniqueness detection, and quality/annotation issue identification.
Install it if you work with computer vision datasets and need systematic quality and similarity analysis.
Provides the database backend for FiftyOne, a computer vision framework for managing and analyzing image and video datasets.
BoxMOT provides pluggable multi-object tracking modules that work with bounding box detections from any model, supporting both axis-aligned and oriented bounding boxes through a unified Python API and CLI.
nnU-Net is a semantic segmentation framework that automatically configures U-Net variants based on dataset characteristics and provides end-to-end workflows for preprocessing, training, model selection, and inference on 2D and 3D images.