Packages
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
Albucore provides optimized atomic image processing functions that automatically select the fastest implementation (NumPy, OpenCV, NumKong, StringZilla, or PyTorch) based on input characteristics for efficient uint8 and float32 image manipulation.
Install it if you are building image augmentation or preprocessing pipelines and want to avoid manual optimization.
kornia-rs provides low-level computer vision operations—image I/O, resizing, color conversion, and video capture—implemented in Rust with Python bindings for efficient, thread-safe processing.
SeleniumBase is a browser automation framework that combines Selenium WebDriver with pytest integration, CDP mode for bot-detection bypass, and tools for web testing, scraping, and crawling.
However, the 60 runtime dependencies are substantial—evaluate whether you need the full feature set or if a lighter alternative suits your use case.
Python client library for Google Earth Engine, enabling programmatic access to satellite imagery and geospatial datasets for remote sensing and environmental analysis.
The main gotcha is that authentication and an Earth Engine project are prerequisites; also, dynamic class loading means you'll rely on the official API Reference.
PyIQA provides a PyTorch-based toolbox for computing image quality assessment metrics, supporting both full-reference and no-reference methods with GPU acceleration and calibration against official implementations.
PyCAT-Napari is a napari-based desktop application for detecting, measuring, and analyzing biomolecular condensates in fluorescence and brightfield microscopy images, with support for 2D, Z-stack, and time-series data.
FiftyOne is a Python framework for building, visualizing, and evaluating computer vision datasets and models, with integrated labeling, model evaluation, and data quality tools.
A Python wrapper for the QOI (Quite OK Image) lossless image format that encodes and decodes images as numpy arrays, offering faster compression and decompression than PNG with comparable file sizes.
ETA is an extensible computer vision and machine learning analytics toolkit that provides core utilities for working with images, videos, embeddings, and ML inference pipelines, along with a CLI for building and running analytics workflows.
FiftyOne Brain provides AI/ML capabilities for analyzing and manipulating datasets and models, including visual similarity search, text-based querying, sample uniqueness detection, and quality/annotation issue identification.
Install it if you work with computer vision datasets and need systematic quality and similarity analysis.
Provides a catalog of pre-defined Pydantic schemas for extracting structured data from images, videos, and documents using Vision Language Models.
Computes eight standard image similarity metrics (RMSE, PSNR, SSIM, FSIM, ISSM, SRE, SAM, UIQ) to quantify how similar two images are, via Python API or command-line tool.
However, do not rely on it for critical production systems or expect bug fixes—consider it a snapshot tool.
Python SDK for the VLM Run API platform, providing access to vision-language models for image and document processing, chat completions, and CLI-based agent interactions.
Lightly provides self-supervised learning models and loss functions for computer vision, enabling you to train neural networks on unlabeled image data using methods like MoCo, SimCLR, BYOL, and DINO.
Install it if you need to train on unlabeled image data or experiment with SSL methods.
Provides the database backend for FiftyOne, a computer vision framework for managing and analyzing image and video datasets.
BoxMOT provides pluggable multi-object tracking modules that work with bounding box detections from any model, supporting both axis-aligned and oriented bounding boxes through a unified Python API and CLI.
Encodes and decodes image data using the JPEG-LS lossless compression algorithm via the CharLS C++ library, producing compressed byte buffers that can be round-tripped with numpy arrays.
However, verify that the fork remains compatible with your use case and that the 655-day release gap does not signal abandonment; check the repository's recent…
TorchIO reads, preprocesses, augments, and samples 3D medical images for deep learning with PyTorch, offering both standard computer vision transforms and domain-specific medical imaging operations like MRI artifact simulation.
Install it if you're building medical imaging models with PyTorch.
Ouster SDK provides Python interfaces to configure, record, and process data from Ouster lidar sensors, including point cloud projection, visualization, and multi-beam flash lidar data handling.