Subcategories
Packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
SymPy is a Python library for symbolic mathematics, performing algebraic manipulation, calculus, equation solving, and mathematical expression simplification without numerical approximation.
Cloudpickle extends Python's standard pickle module to serialize lambda functions, interactively-defined functions and classes, and other constructs that the default pickle cannot handle, making it suitable for cluster computing and remote code execution.
Install it if you need to serialize lambda functions, interactively-defined code, or non-standard Python constructs for cluster computing or distributed execution.
PyTorch provides GPU-accelerated tensor computation and automatic differentiation for building and training deep neural networks in Python.
Provides type stubs for pandas, enabling static type checkers like mypy and pyright to catch type errors in pandas code before runtime.
onnxruntime loads and executes Open Neural Network Exchange (ONNX) models with a focus on inference performance across CPUs and accelerators.
Install it if you have ONNX models to run in production or development.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Polars is a DataFrame query engine written in Rust that executes analytical queries with multi-threaded, vectorized performance, supporting both lazy and eager evaluation modes.
opencv-python provides pre-built CPU-only OpenCV bindings for Python, enabling computer vision tasks like image processing, object detection, and video analysis without requiring separate OpenCV installation.
Jupyter Notebook is a web-based interactive computing environment that lets you create and run code, visualizations, and documentation in a single browser-based interface, supporting multiple programming languages through pluggable kernels.
Install it if you need an interactive notebook environment for exploration, education, or reproducible research.
DuckDB is an in-process SQL database engine that executes analytical queries on local and remote data files (CSV, Parquet, JSON) without requiring a separate server or external dependencies.
Polars-runtime-32 is a compiled runtime component for an analytical query engine that executes DataFrame queries with multi-threaded, vectorized performance and supports both lazy and eager evaluation modes.
JupyterLab is an extensible web-based interactive computing environment that provides notebooks, terminals, text editors, and file browsers in a unified interface for reproducible computing workflows.
Install it if you need an interactive computing environment for research, data science, or exploratory development.
Provides NVIDIA's NCCL runtime library for GPU collective communication operations including all-reduce, all-gather, reduce, broadcast, and reduce-scatter.
Install only if your system has compatible NVIDIA GPUs and CUDA 12 already installed.
Provides low-level Python bindings for CUDA host APIs, offering 1:1 access to CUDA functionality from Python code.
OpenCV Python bindings for computer vision tasks, packaged without GUI dependencies for server and headless environments.
Install it if you need computer vision in a server or containerized environment where GUI dependencies are unnecessary.
h5py reads and writes HDF5 files from Python, providing both low-level API access and high-level interfaces that map HDF5 datasets and groups to NumPy arrays and Python data structures.
Install it if you work with scientific data, need portable binary storage, or must interoperate with HDF5 files from other tools or languages.
statsmodels provides statistical models, inference methods, and descriptive statistics for Python, complementing scipy with regression, time series, discrete choice, survival analysis, and multivariate methods.
Install it if you need publication-quality statistical models, hypothesis tests, or time series analysis beyond what scipy or pandas provide.
Provides NVIDIA CUBLAS native runtime libraries for GPU-accelerated linear algebra operations in CUDA-enabled environments.
Provides NVIDIA CUDA NVRTC (NVIDIA Runtime Compilation) native runtime libraries for Python, enabling runtime compilation of CUDA kernels on x86_64 Linux, ARM64 Linux, and Windows platforms.
Provides NVIDIA's JIT LTO compiler library for Python, enabling just-in-time and link-time optimization compilation functionality for CUDA-based applications.
Provides cuDNN runtime libraries for GPU-accelerated deep neural network primitives, requiring CUDA 13 and nvidia-cublas as dependencies.
Install only if you have CUDA 13 and nvidia-cublas available; otherwise, installation will not resolve the underlying GPU dependencies.
Provides NVIDIA's Collective Communication Library (NCCL) runtime for GPU-accelerated collective operations like all-reduce, all-gather, reduce, broadcast, and reduce-scatter across multiple GPUs using PCIe, NVLink, NVswitch, InfiniBand, or TCP/IP.
However, verify that your framework (PyTorch, TensorFlow, etc.) declares it as a dependency rather than installing it standalone.
Provides NVIDIA's CUDA library for high-performance sparse matrix-matrix multiplication on GPUs with structured sparsity, supporting mixed-precision computation across multiple data types.
However, verify NVIDIA's proprietary license terms for your use case, confirm your GPU architecture is supported (SM 8.0+), and ensure CUDA 13 is installed and…
NVSHMEM provides a global address space for GPU cluster communication, enabling fine-grained GPU-initiated and CPU-initiated operations across multiple GPUs' memory via a parallel programming interface based on OpenSHMEM.
However, the unclear license status and lack of public documentation are concerns—verify licensing terms and review NVIDIA's CUDA Zone documentation before committing…
Provides NVIDIA CUSPARSE native runtime libraries for GPU-accelerated sparse matrix operations in CUDA applications.
Install only if you have NVIDIA GPU hardware and the CUDA toolkit already set up; this is a runtime library, not a standalone tool.
A meta-package that installs NVIDIA CUDA Toolkit components via pip, allowing selective installation of GPU-accelerated libraries, compilers, and runtime tools through package extras.
However, verify license terms beforehand and confirm your GPU and driver support the toolkit version—this is not a substitute for proper CUDA environment setup on…
Provides NVIDIA CUFFT native runtime libraries for GPU-accelerated Fast Fourier Transform computations on CUDA-enabled hardware.
Provides CUDA solver native runtime libraries for GPU-accelerated linear algebra operations on NVIDIA hardware.
Provides NVIDIA CURAND native runtime libraries for GPU-accelerated random number generation in Python on Linux, Windows, and aarch64 platforms.
Provides NVIDIA CUDA Runtime native libraries for GPU-accelerated computing on Windows and Linux systems.
Provides CUDA profiling runtime libraries that enable third-party tools to access GPU profiling APIs on NVIDIA hardware.
However, it requires an NVIDIA GPU and CUDA environment, and the unclear license should be confirmed before commercial deployment.
Provides Python bindings for NVIDIA's cuFile GPUDirect storage access libraries, enabling direct GPU-to-storage I/O without CPU involvement.
However, verify the unclear license terms before use in production, and confirm that your system has the required cuFile runtime libraries installed.
Provides a Python API for annotating events and code ranges in applications to capture performance profiling data visible in NVIDIA's Visual Profiler.
Install only if you have CUDA and Visual Profiler available; otherwise it provides no value.
Patsy converts R-style statistical formulas into design matrices for Python, enabling compact specification of linear models and model components without manual matrix construction.
Install it if you need R-style formula notation for design matrices or are porting R code.