Packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
Snowflake SQLAlchemy is a SQLAlchemy dialect that enables SQLAlchemy applications to connect to and query Snowflake databases using standard SQLAlchemy ORM and Core APIs.
Install it if you are building a Python application that needs to connect to Snowflake and prefer SQLAlchemy's abstraction layer over raw SQL or the connector API.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
fastavro reads and writes Apache Avro files with C-extension performance, supporting schemaless operations, multiple codecs, and schema resolution.
Install it if you work with Avro files or need fast serialization of structured data.
agate is a Python data analysis library designed for readable, human-friendly code that handles tabular data manipulation, filtering, aggregation, and analysis without requiring numpy or pandas.
Streamlit transforms Python scripts into interactive web applications with minimal code, enabling rapid development of data dashboards, reports, and chat interfaces without requiring web development expertise.
Read and write TIFF files and TIFF-like formats (BigTIFF, OME-TIFF, GeoTIFF, and many bioimaging formats) with NumPy array integration and support for multiple compression schemes.
Install it if you work with TIFF files, microscopy data, or need to convert between TIFF variants.
Leather is a lightweight Python charting library for quick, no-frills data visualization. It generates charts without requiring perfect styling or extensive configuration.
natsort provides natural sorting for strings containing numbers, ordering them by numeric value rather than lexicographic order—so '10' comes after '9' instead of after '1'.
Install it if you need to sort strings containing numbers in a human-readable way.
A native TCP driver for ClickHouse that executes queries and returns results, supporting both a direct client interface and Python DB API 2.0 specification.
FastF1 provides Python access to Formula 1 timing data, telemetry, results, and schedules through an Ergast-compatible API, returning data as extended Pandas DataFrames with F1-specific analysis functions and Matplotlib visualization integration.
Install it if you work with Formula 1 data; skip it if you have no need for F1-specific APIs or telemetry.
Extracts original and updated publication dates from web pages by parsing HTML markup, metadata, and text content, with both Python API and command-line interfaces.
Install it if you need reliable date extraction from web pages; the fast mode offers good speed and the extensive mode provides high recall when accuracy matters most.
PyPika is a Python API for building SQL queries programmatically using a builder pattern, eliminating string formatting and concatenation while maintaining flexibility for complex queries.
Install it if you need to build queries dynamically or support multiple SQL dialects from Python code.
Trafilatura extracts main text, metadata, and structured content from web pages and HTML, converting raw HTML into clean, usable data in multiple output formats.
Validates, normalizes, filters, and samples URLs for web crawling and document collection, removing spam, trackers, and low-value pages while respecting language and content-type constraints.
Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.
Torchmetrics provides a collection of PyTorch metrics implementations with automatic batch accumulation and multi-device synchronization, designed for distributed training workflows.
Bokeh is an interactive visualization library that creates browser-based plots, dashboards, and data applications from Python code, with support for large and streaming datasets.
Install it if you need browser-based interactivity.
PyTorch Lightning wraps PyTorch training code to separate research logic from engineering boilerplate, enabling distributed training across GPUs, TPUs, and CPUs without code changes.
Swifter applies functions to DataFrames and Series using automatic vectorization or parallel processing to speed up operations beyond standard apply.
However, proceed with caution: maintenance is dormant (last release 2023-07-31), license status is unclear, and Python version support is unspecified.
Cython-accelerated implementation of the toolz functional utilities library, providing faster performance for operations on iterables, functions, and dictionaries while maintaining the same API.
Gensim is a Python library for topic modeling, document indexing, and similarity retrieval on large text corpora, using algorithms like LDA, LSA, word2vec, and others.
However, the aging maintenance status (300 days since last release) suggests the project is in steady-state rather than actively developed; evaluate whether its…
Provides type annotations and runtime type-checking for array shape and dtype across JAX, PyTorch, NumPy, MLX, and TensorFlow, with no JAX dependency required.
Prophet forecasts time series data using an additive model that combines non-linear trends with seasonal patterns (yearly, weekly, daily) and holiday effects, handling missing data and outliers robustly.
Install it if your data has clear seasonal patterns and you want a library that requires minimal tuning; skip it if you need ultra-low latency inference or…
CmdStanPy provides a pure-Python interface to the Stan probabilistic programming language, enabling you to compile Stan models and run Bayesian inference algorithms without direct C++ interaction.
Install it if you need to run Stan models from Python for Bayesian inference, statistical modeling, or probabilistic programming.
Parsimonious is a pure-Python PEG (parsing expression grammar) parser that builds abstract syntax trees from text according to grammar rules you define, with no external lexer or parser generator required.
Provides probabilistic data structures (MinHash, HyperLogLog, and related indexes) for fast similarity estimation and cardinality counting on large datasets with minimal memory overhead.
Lightning is a framework that organizes PyTorch code to automate distributed training infrastructure—handling backpropagation, mixed precision, multi-GPU, and multi-node setups—while keeping your model logic unchanged.
Python binding to CRFsuite for conditional random field sequence labeling and structured prediction.
Docling Core defines the foundational DoclingDocument data model and provides APIs for serialization, chunking, and profiling of structured document data for generative AI applications.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Provides Python access to Snowflake entity metadata and resource management, allowing you to create, delete, and modify Snowflake resources programmatically.
Install it if you need to automate Snowflake object lifecycle operations from Python; skip it if you only need to query data (use snowflake-connector-python directly…
vLLM is a high-throughput inference and serving engine for large language models, offering optimized memory management, continuous batching, and support for multiple hardware platforms and model architectures.
Install it if you're building an LLM application, API service, or batch inference pipeline; skip it if you only need simple single-model inference without serving…
DDSketch computes quantiles (percentiles) of streaming or batch data with guaranteed relative error bounds, and supports merging sketches from distributed systems into a single combined sketch.
Detects sentence boundaries in text using rule-based heuristics, splitting paragraphs into individual sentences while handling edge cases like abbreviations and decimal points.
However, do not install if you need active maintenance, multi-language support, or assurance of ongoing bug fixes.
Provides common utility methods for building probable parsers, a parsing approach used in information extraction and data cleaning tasks.