Packages
Modin is a drop-in replacement for pandas that distributes DataFrame operations across multiple CPU cores, enabling faster execution on larger datasets that would otherwise exhaust memory or run slowly with single-threaded pandas.
However, the 316-day gap since the last release suggests slower maintenance cadence; verify that your specific pandas operations are fully supported before committing…
Awkward Array provides NumPy-like operations on nested, variable-sized data structures—lists, records, mixed types, and missing values—with compiled performance and dynamic typing.
Install it if your workflow involves JSON-like structures, ragged arrays, or hierarchical data that NumPy alone cannot handle efficiently.
PyVista provides a NumPy-native API for 3D visualization and mesh analysis, with dataset structures and filters for points, surfaces, and volumes, plus a unified plotting framework for notebooks, scripts, CI, and applications.
Install it if you need 3D mesh visualization, point-cloud analysis, or volumetric data exploration.
A command-line toolkit that simplifies common bioinformatics tasks like fetching sequences, converting formats, aligning genomes, and querying taxonomic data through composable commands.
Visions defines and detects semantic data types across pandas, numpy, spark, and Python sequences, automatically inferring the most appropriate type even when data has been transformed or corrupted.
Chonkie splits text into semantically meaningful chunks for RAG pipelines, offering multiple chunking strategies (recursive, semantic, token-based, code-aware) plus refinement and embedding integration.
Reads and writes GRIB and BUFR meteorological data files via Python bindings to the ECMWF ecCodes C library.
Install it if you work with GRIB or BUFR files.
Kedro-Viz is an interactive web-based visualization tool for Kedro data science pipelines that displays pipeline structure, parameters, and metadata in a searchable, filterable interface.
Equinox provides neural network and model building on top of JAX with PyTorch-like syntax, plus PyTree manipulation, filtered transformations, and runtime error handling—all while remaining fully compatible with core JAX operations.
Provides compiled C++ kernels and extensions that accelerate operations on nested, variable-sized data structures with performance comparable to NumPy.
Phi_K computes a correlation coefficient that works consistently across categorical, ordinal, and interval variables, capturing non-linear dependencies while reverting to Pearson correlation for bivariate normal data.
However, the aging maintenance status (no release in 393 days) suggests checking whether active development continues before relying on it for production systems.
Provides streaming algorithms (sketches) for approximate answers to big-data queries like cardinality estimation, quantiles, and frequent items with proven error bounds and orders-of-magnitude speed gains.
Spark NLP provides distributed natural language processing on Apache Spark, offering pretrained pipelines and models for tokenization, named entity recognition, sentiment analysis, machine translation, and embeddings across multiple languages.
Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.
Splink performs probabilistic record linkage and deduplication, matching records across datasets that lack unique identifiers by comparing multiple columns and assigning match probabilities.
Uproot reads and writes ROOT files (the data format used in high-energy physics) directly in Python using NumPy, without requiring the C++ ROOT library.
Install it if you work with ROOT files or need to integrate them into Python data pipelines.
Mlxtend provides ensemble methods, feature selection, visualization utilities, and frequent pattern mining algorithms for machine learning workflows.
Install it if you need these specific capabilities.
Stanza is a Python NLP library that runs accurate natural language processing tools on 60+ languages, including tokenization, part-of-speech tagging, dependency parsing, and named entity recognition, with optional access to Java Stanford CoreNLP.
Install it if you need dependency parsing, NER, or POS tagging across many languages or in biomedical domains; skip it only if you need real-time performance on…
Implements OWL2 RL Profile and RDFS inference on top of rdflib using forward-chaining rules to derive new triples from RDF graphs.
Koalas implements the pandas DataFrame API on top of Apache Spark, letting you write pandas-like code that runs on distributed Spark clusters instead of a single machine.
Python client library for IBM Watson Machine Learning service, enabling model training, testing, and deployment as APIs on IBM Cloud and IBM Cloud Pak for Data.
A lightweight, pluggable HTTP proxy server framework that handles forward proxying, reverse proxying, and TLS interception with support for custom plugins and asyncio-based concurrency.
Weave is a toolkit for tracing, logging, and evaluating generative AI applications, letting you instrument functions to capture inputs, outputs, and execution traces for debugging and analysis.
Snowflake ML Python provides SDKs and infrastructure to build, train, manage, and deploy machine learning models directly within Snowflake, covering data preprocessing, feature engineering, model development, experiment tracking, and model registry.
eccodeslib provides Python bindings to decode and encode meteorological data in GRIB, BUFR, and WMO GTS formats, wrapping the ECMWF ecCodes C library.
boost-histogram provides Python bindings to Boost::Histogram, a C++14 library for fast histogram creation and manipulation with support for multiple axis types, storage modes, and advanced indexing.
Connects Python notebooks and Spark jobs in Microsoft Fabric to Power BI datasets, enabling data scientists to query semantic models, augment data with Power BI measures, and propagate semantic information across analysis workflows.
However, the proprietary license and Beta status mean you should review Microsoft's terms and test thoroughly before production use.
Computes technical analysis indicators (volume, volatility, trend) from financial time series data for feature engineering in trading and financial analysis workflows.
However, verify Python version compatibility and build requirements before committing to production use.
EdgarTools parses SEC EDGAR filings into typed Python objects and pandas DataFrames, extracting financial statements, insider trades, fund holdings, and other filing types with a consistent API.
The main gotcha is that you must provide an email to the SEC with every request, but that is a documented requirement.
Pymatgen provides Python classes and analysis tools for materials science workflows, including structure representation, file I/O for computational chemistry formats (VASP, ABINIT, Gaussian, CIF), phase diagrams, electronic structure analysis, and integration with the Materials Project REST API.
Install it if you work with crystal structures, computational chemistry outputs, or materials databases; skip it if your work does not involve materials analysis.
cfgrib maps GRIB meteorological data files to the NetCDF Common Data Model following CF Conventions, exposing them as labeled multi-dimensional datasets via an engine interface.
Computes Short Term Objective Intelligibility (STOI) measures for speech signals, quantifying how intelligible degraded audio is compared to a clean reference.
However, verify compatibility with your Python version and confirm the codebase meets your performance needs, since the last release was 959 days ago.
Provides over 150 technical analysis indicators and 60 candlestick patterns for financial data, optimized with numba and numpy, and integrated as a pandas DataFrame extension.
However, the Beta status, 334-day maintenance gap, and unclear license terms warrant caution—verify the license for your use case and be prepared for potential…
Reads, writes, and analyzes SVG Path objects and Bézier curves, providing geometric tools to transform, intersect, and measure path elements.
Rounds numbers by significant figures, decimal places, or uncertainty, and formats them in multiple scientific and publication styles with results that match expected mathematical behavior.
SimpleITK provides a simplified Python interface to the Insight Toolkit (ITK) for performing image segmentation, registration, and general filtering operations on 2D, 3D, and 4D images.
mplfinance provides matplotlib-based visualization for financial market data, enabling candlestick charts, OHLC plots, and technical analysis overlays from pandas DataFrames.
However, do not rely on it for active bug fixes or new features—test compatibility with your matplotlib and pandas versions before production use, and consider it a…
Implements the Leiden community detection algorithm for graphs, exposing a C++ implementation to Python via igraph for partitioning networks into communities using multiple optimization methods.
Splits Indo-European text into sentences and words using rule-based segmentation and tokenization, with command-line tools for batch processing.
Implements the W3C PROV Data Model in Python, enabling creation, import, and export of provenance documents in multiple formats including PROV-O, PROV-XML, PROV-JSON, and PROV-JSONLD.
Install it if you need to work with provenance data in a standards-compliant way or integrate with systems that consume PROV documents.
Hist provides a user-friendly interface for creating, filling, and analyzing histograms, built on top of boost-histogram with support for named axes, advanced indexing, and integrated plotting.