Subcategories
Packages
Jupyter is a metapackage that installs the core Jupyter components—Notebook, JupyterLab, IPython Kernel, and related tools—in a single command for interactive computing and data science workflows.
ppft executes Python functions in parallel across multiple processors or networked computers, bypassing the GIL by using separate processes and inter-process communication.
Optuna is a hyperparameter optimization framework that automates the search for optimal hyperparameter values in machine learning models using a define-by-run API and state-of-the-art sampling algorithms.
Pooch downloads files from HTTP, FTP, and data repositories (Zenodo, figshare), caches them locally, and verifies integrity via hash checking—eliminating manual file management and urllib boilerplate.
Install it if you need to download, cache, and verify data files in any Python project—it eliminates boilerplate and adds reproducibility with minimal overhead.
Provides re-sampling techniques to address class imbalance in machine learning datasets, integrating with scikit-learn for preprocessing imbalanced data before model training.
Kubeflow Pipelines is a Python SDK for defining, deploying, and managing machine learning workflows as containerized task graphs on Kubernetes clusters.
Install only if you already have or plan to set up a Kubeflow infrastructure; it is not a standalone ML framework.
Timm provides a large collection of pretrained PyTorch image models—vision transformers, CNNs, and hybrid architectures—with utilities for loading weights, training, and inference.
Install it if you need pretrained models, model zoo access, or a training framework for vision tasks.
Provides Python bindings for NVIDIA's cuFile GPUDirect storage access libraries, enabling direct GPU-to-storage I/O without CPU involvement for CUDA 12 environments.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
tritonclient is a Python client library for communicating with Triton Inference Server over gRPC or HTTP, enabling you to send inference requests to remote model-serving deployments.
Install it if you are using Triton for model serving and need to query it from Python.
Trimesh loads, manipulates, and analyzes triangular mesh geometry in pure Python, with emphasis on watertight surfaces and a single hard dependency on numpy.
RQ is a Python library for queueing jobs and processing them in the background with workers, backed by Redis or Valkey. It handles job scheduling, prioritization, retries, and webhooks with a simple API.
Bokeh is an interactive visualization library that creates browser-based plots, dashboards, and data applications from Python code, with support for large and streaming datasets.
Install it if you need browser-based interactivity.
RDKit is a chemoinformatics toolkit providing C++ and Python libraries for molecular structure manipulation, analysis, and machine-learning workflows on chemical data.
Install it if you work with molecular structures, chemical data, or computational chemistry.
Calls any Gradio app as a Python API with a simple Client interface, handling authentication, file uploads, and response parsing automatically.
A build backend that uses CMake to compile Python extension modules and packages, replacing setuptools-based build systems with a modern, standards-compliant approach.
Performs high-quality sample-rate conversion (resampling) for audio signals, supporting both one-shot and streaming modes via a Python wrapper around libsoxr.
Dask Expressions provides query optimization for Dask DataFrames by encoding operations as an expression tree that is optimized before execution, replacing the earlier Dask DataFrame implementation.
Install it if you are using Dask DataFrames; it is the recommended path forward.
Bottleneck provides fast NumPy array functions written in C, accelerating operations like nanmean, nansum, moving window calculations, and ranking on arrays with NaN values.
cuda-python is a metapackage providing Pythonic access to NVIDIA's CUDA platform, bundling low-level bindings, runtime abstractions, and utilities for GPU-accelerated computing from Python.
However, verify that the proprietary NVIDIA license aligns with your project's requirements before committing to it in production or redistributed code.
TensorFlow Estimator provides a high-level API for building and training machine learning models, encapsulating training, evaluation, prediction, and model export workflows.
opencv-contrib-python provides Python bindings to OpenCV's full library, including contrib modules for advanced computer vision tasks like feature detection, image processing, and machine learning on images.
Pandera provides a flexible API for validating dataframe-like objects using declarative schemas with type checking and custom validation rules.
Cython-accelerated implementation of the toolz functional utilities library, providing faster performance for operations on iterables, functions, and dictionaries while maintaining the same API.
Ultralytics YOLO provides a unified framework for training and deploying computer vision models for object detection, instance segmentation, pose estimation, image classification, and semantic segmentation tasks.
Install it if you need a production-ready YOLO implementation and can accept the AGPL-3.0 license (or obtain a commercial license).
GluonTS provides deep learning models for probabilistic time series forecasting, built on PyTorch, enabling you to train and deploy models that generate probability distributions over future values rather than point estimates.
Install it if your use case requires neural time series models; skip it if you need only classical statistical forecasting or simple point predictions.
Distributed provides a scheduler and runtime for parallel and distributed computation using Dask, enabling you to scale Python workloads across multiple machines or cores.
LanceDB is a Python SDK for connecting to and querying LanceDB vector databases, supporting search operations on embedded data with filtering and result limits.
NVSHMEM provides a global address space for GPU cluster communication, enabling fine-grained GPU and CPU-initiated operations across multiple GPU memories using OpenSHMEM-based primitives.
However, verify the unclear license terms and confirm Windows support is not actually available despite classifier claims.
UMAP reduces high-dimensional data to lower dimensions for visualization and general non-linear dimensionality reduction, using a manifold-learning approach that preserves both local and global structure.
Install it if you need dimension reduction with better global structure preservation, or as a general-purpose manifold learning tool for exploratory analysis and…
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
Provides trading calendars for more than 50 security exchanges worldwide, allowing you to query trading sessions, minutes, and schedules with timezone-aware timestamps and holiday handling.
Install it if you work with multiple exchanges or need reliable session/minute-level time alignment.
PyNNDescent builds approximate nearest neighbor search indexes using nearest neighbor descent algorithms, supporting a wide variety of distance metrics for fast k-neighbor queries on high-dimensional data.
OR-Tools provides constraint programming, linear and mixed-integer programming, vehicle routing, and graph algorithm solvers developed at Google for operations research problems.
Provides enhanced HTTPS support for Python's httplib and urllib2 (or http.client and urllib in Python 3) by wrapping PyOpenSSL to enable full SSL peer verification using certificate validation.
Provides Google Cloud Storage filesystem support for TensorFlow, enabling direct reading and writing of data from GCS buckets within TensorFlow pipelines without local downloads.
Install only if your TensorFlow version matches the compatibility table (0.37.1 requires TensorFlow 2.16.x).
Profiles PyTorch models by counting Multiply-Accumulate Operations (MACs) and parameters in a single forward pass, with built-in rules for common layer types and support for custom counting rules.
Install it if you need to profile PyTorch model efficiency.
Python binding to CRFsuite for conditional random field sequence labeling and structured prediction.
Provides Python interfaces for writing high-performance CUDA kernels using CUTLASS DSL concepts without C++ expertise, targeting NVIDIA Tensor Cores on Ampere, Hopper, and Blackwell architectures.
However, verify the unclear license terms before production use, and be aware the package is in public beta—expect potential API changes before summer 2026 graduation.
Pint defines, operates on, and converts between physical quantities—combining numerical values with units of measurement—and supports arithmetic operations and unit conversions across a comprehensive built-in unit library.
Install it if you work with physical quantities, scientific data, or any calculation where mixing units is a concern.