Packages
usaddress parses unstructured US address strings into labeled components using a probabilistic model trained on conditional random fields, handling ambiguous cases where rule-based parsers fail.
Provides memory-efficient trie data structures for fast string lookups and prefix searches, using up to 50x-100x less memory than standard Python dicts while maintaining comparable lookup speed.
PySTAC is a Python library for reading, writing, and working with SpatioTemporal Asset Catalog (STAC) specifications, enabling standardized organization and discovery of geospatial and temporal data.
Install it if you work with satellite imagery, earth observation data, or need to organize and catalog geospatial assets according to the STAC standard.
PyGraphviz provides a Python interface to Graphviz, enabling you to create, edit, read, write, and draw graphs using Graphviz's layout algorithms and visualization engine.
Tempo provides time series operations on Spark DataFrames, including AS OF joins, rolling statistics, lagged feature generation, and Delta Lake optimization for time-partitioned data.
Install it if you work with time-indexed data; skip it if you do not need time series transformations.
PySTAC Core provides the foundational API for reading, writing, and working with SpatioTemporal Asset Catalog (STAC) metadata in Python, handling catalog structure, assets, and item management without extension implementations.
Extends PySTAC with the View Geometry Extension, adding fields to describe observation angles and sun geometry for geospatial imagery data.
Install only if your workflow actually requires view geometry fields; base PySTAC alone is sufficient for simpler use cases.
Adds version tracking and deprecation status fields to PySTAC catalogs, enabling STAC items and collections to record version history and link to predecessor and successor versions.
Extends PySTAC to support the Xarray Assets Extension specification, enabling standardized metadata for describing how STAC assets can be opened with xarray, including storage options and engine configuration.
Provides PySTAC extension support for describing categorical data and classification schemes in raster and vector data, including class names, descriptions, color hints, and regional mappings.
Extends PySTAC to add cloud storage metadata fields (platform, region, tier, requester-pays) to STAC assets via the Storage Extension specification.
Install it if you are building or extending STAC catalogs and need to document cloud storage attributes for your assets.
Extends PySTAC to add scientific publication metadata fields—DOIs, citations, and bibliographic references—to STAC items and collections.
Extends PySTAC with fields and validation for Synthetic-Aperture Radar (SAR) data, enabling description of instrument mode, frequency band, polarizations, product type, and observation direction in STAC catalogs.
Install it if you are working with SAR data in a PySTAC-based system and need standardized metadata fields for radar imagery.
Extends PySTAC to add File Info Extension support, enabling description of file metadata like byte order, checksum, header size, and data type for catalog assets.
Install it if you are building or extending STAC catalogs and need to attach standardized file metadata to assets.
Adds MGRS (Military Grid Reference System) extension support to PySTAC, enabling description of geospatial data using UTM zones, latitude bands, and grid square identifiers.
Install it if you need to catalog geospatial data using MGRS grid references within a PySTAC workflow; skip it if your data does not require military grid coordinate…
Extends PySTAC to describe gridded data products by specifying grid codes like MGRS or MODIS tiling schemes that data corresponds to.
Install it if you are working with PySTAC and need to describe gridded data products or integrate grid metadata into STAC catalogs.
Extends PySTAC to add Electro-Optical extension support for describing sensor data, spectral bands, cloud cover, and snow cover in STAC catalogs.
Extends PySTAC to support the Render Extension specification, enabling description of visualization parameters for raster assets including rescaling, color maps, expressions, and tile matrix set configurations.
Extends PySTAC to support the Point Cloud Extension specification, enabling description and cataloging of point cloud datasets with metadata for point count, encoding, density, schemas, and statistics.
Extends PySTAC to support the Datacube Extension specification, enabling representation of multi-dimensional datasets with dimension types, extents, values, and reference systems.
Install only if your use case requires explicit datacube dimension metadata; pystac-core alone may suffice for simpler catalogs.
Extends PySTAC to handle satellite-specific metadata fields like orbit state vectors, relative orbit numbers, and platform identifiers through the Satellite Extension specification.
Unified Python API for Snowflake workloads, providing access to data engineering, Snowpark, Snowpark ML, and client application resources through a single namespace package.
SimSIMD provides SIMD-optimized kernels for computing vector distances, dot-products, and similarity measures across multiple data types and precisions, with support for spatial, probabilistic, and bit-level operations.
Install it if vector similarity or distance computation is a measurable bottleneck in your application.
Panel is a Python framework for building interactive data applications, dashboards, and web apps with widgets, plots, and tables that can be deployed as web services, notebooks, or static exports.
Install it if you need to turn Python data work into interactive web apps or dashboards without learning JavaScript or web frameworks.
Provides temporary backward compatibility for code that imports the old unrelated `snowflake` package, allowing it to read from `/etc/snowflake` or an alternative path; this is a migration bridge, not a primary tool.
igraph provides a Python interface to a high-performance C graph library for constructing, analyzing, and visualizing networks and complex graphs.
Parses and compiles Excel formulas and workbooks into executable Python code without requiring Excel itself.
The copyleft license (EUPL 1.1+) requires careful review if you plan to distribute derivative works or integrate into proprietary software; internal use is generally…
Schedula is a flow-based programming library that automatically manages control flow by building a directed acyclic graph (DAG) of functions and their data dependencies, then selects and executes the optimal path to compute requested outputs from given inputs.
Install it if you need automatic dataflow scheduling and can accept EUPL 1.1+ copyleft obligations.
Generates STAC (SpatioTemporal Asset Catalog) items and collections for Met Office UK deterministic weather forecast data, enabling standardized metadata and discovery of weather model outputs.
However, the aging maintenance status (no release in 198 days) means you should verify that it meets your current needs before committing; check whether it handles…
VTK is a 3D graphics, image processing, and visualization toolkit that provides algorithms for surface reconstruction, volume rendering, and advanced rendering techniques.
RecordLinkage identifies and matches records within or across datasets using indexing, comparison, and classification algorithms, supporting both deduplication and cross-dataset linking tasks.
However, avoid it for production systems requiring active support or if you need compatibility guarantees with very recent versions of numpy, pandas, or scikit-learn.
AKShare fetches financial market data (stocks, options, futures, bonds, funds) from multiple Chinese and international sources via a unified Python API, with built-in interface search and metadata lookup.
StringZilla provides SIMD and SWAR-accelerated string operations including substring search, hashing, edit distances, sorting, and segmentation for Python, with no runtime dependencies.
Install only if your Python version is 3.10 or later.
Detects which language a text is written in, supporting 75 languages with high accuracy on both short snippets and full sentences using compiled Rust bindings.
Provides data structures for representing and manipulating temporal segments with labels, designed for tasks like speaker diarization and time-based annotation workflows.
However, verify the license status before use, and note that maintenance appears to be in a slower phase (332 days since last release)—acceptable for stable…
Official Python SDK for IBM watsonx.ai that provides a unified interface to foundation models, AutoAI experiments, retrieval-augmented generation, model tuning, and deployment across the watsonx.ai platform.
Install it if you are building or deploying AI models on the watsonx.ai platform; skip it if you are not using that specific IBM service.
Detects and replaces personally identifiable information (names, emails, phone numbers, credit cards, dates of birth, social security numbers, and more) in free text with anonymized placeholders.
However, maintenance is dormant, so if you need active bug fixes or feature development, evaluate whether the package's current detector coverage meets your needs.
SQLAlchemy dialect for ClickHouse that enables ORM-style database access via native TCP, async TCP, or HTTP interfaces.
However, you must manually install clickhouse-driver, asynch, or requests depending on your transport choice—verify this requirement before adding it to your project.
Autovizwidget automatically generates interactive visualizations for pandas dataframes in Jupyter notebooks, integrating with Sparkmagic's remote Spark cluster execution.
However, the aging maintenance status (403 days since last release) and lack of recent updates suggest you should verify it works with your current Jupyter and…
Provides native Python and C++ bindings to parse, read, and manipulate mmCIF (macromolecular Crystallographic Information File) data structures and dictionaries.