Packages
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Provides Python utilities for writing pandoc filters that transform document ASTs by reading JSON from stdin, modifying the structure, and writing JSON to stdout.
Converts Unicode text to ASCII-safe representations by transliterating non-ASCII characters to their closest ASCII equivalents, useful for generating URL slugs, legacy system integration, and keyboard-friendly identifiers.
Install it if you need lossy Unicode-to-ASCII conversion and can accept GPLv2+ copyleft terms.
Pymdown Extensions provides a collection of additional extensions for Python Markdown that add syntax support for advanced formatting, including tabbed content, superfences with syntax highlighting, emoji, and other markup enhancements.
Install it if you need markdown features beyond the standard library.
Provides emoji and SVG icon extensions for MkDocs Material, enabling insertion of Material, FontAwesome, and Octicons into Markdown using emoji syntax.
Represents and manipulates IPv4, IPv6, MAC addresses and related network objects; supports CIDR notation, subnetting, set operations, IANA lookups, and DNS reverse generation.
Validates, normalizes, filters, and samples URLs for web crawling and document collection, removing spam, trackers, and low-value pages while respecting language and content-type constraints.
Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.
jusText removes boilerplate content (navigation, headers, footers) from HTML pages while preserving text with full sentences, useful for extracting clean content for linguistic analysis and web corpora.
However, maintenance is aging (535 days since last release), so it is best suited for established use cases where the algorithm is known to work well for your content…
Converts Markdown text to HTML using a fast, complete Python implementation of the Markdown specification.
Install it if you need reliable Markdown rendering without external dependencies.
Pypandoc wraps pandoc, a universal document converter, letting you convert between document formats (Markdown, HTML, DOCX, PDF, and many others) from Python code.
Install it if you need to convert between document formats from Python code.
Extracts encapsulated HTML and plain text content from RTF bodies in .msg email files, reversing the RTF encapsulation that Microsoft Exchange applies.
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
Converts marked-up text between formats including plain text, XHTML, RTF, and PDF, preserving basic structure like paragraphs, headings, lists, and simple tables.
pypandoc_binary wraps pandoc, a universal document converter, and bundles the pandoc binary so you can convert between document formats without a separate pandoc installation.
rjsmin minifies JavaScript code by removing unnecessary characters while preserving functionality, implemented in Python with optional C acceleration for runtime use.
However, consider the 306-day release gap when evaluating long-term support; if you require active maintenance or are starting a new project, evaluate whether a more…
RCSSmin minifies CSS by removing spaces, comments, and unnecessary characters while preserving CSS semantics and supporting common CSS hacks.
However, the 306-day gap since the last release and aging maintenance status suggest limited active development; verify that the package still meets your platform and…
Provides a Python codec to convert between LaTeX-encoded text and Unicode, suitable for processing short text fragments like BibTeX entries or paragraphs rather than full documents.
Parses the Public Suffix List to extract the public and private (registrable) portions of domain names, with support for internationalized domain names and Punycode.
Install it if you need reliable public suffix parsing; the daily-updated built-in list and pure-Python implementation make it a practical choice.
Detects Unicode homoglyphs and mixed-script strings that could be used in spoofing attacks, helping prevent homograph attacks where visually similar characters trick users.
However, be aware the repository is archived and unmaintained—unicode data is current as of 2024-01-30, but you should monitor whether future Unicode standards…
Backports Python 3.6's textwrap module to earlier Python versions, making APIs like shorten() and max_lines available across Python 2.6 and 3.x with improved Unicode handling.
Minifies JavaScript code by removing unnecessary characters while preserving functionality, available as both a Python library and command-line tool.
Wraps and formats text while correctly handling ANSI color and style codes, treating them as invisible to line length calculations the way textwrap cannot.
Install it if you're building terminal UIs or CLI tools with colored output that need proper text wrapping.
A Markdown extension that converts PlantUML diagram syntax into rendered images (PNG, SVG, or text) embedded in HTML documents, using either a local PlantUML binary or a remote server.
Converts markdown documents to plain text by stripping formatting and rendering only the text content.
Converts text to title case with intelligent handling of small words (a, an, the) based on New York Times style rules, plus a command-line utility for batch processing.
Not recommended if you need active bug fixes or compatibility guarantees beyond Python 3.10.
EmPy is a templating system that embeds Python code directly into text documents using customizable markup (default `@`), processing the combined source to produce output with Python expressions, statements, and control structures evaluated inline.
Converts natural language number words into their digit representations across seven languages, and detects ordinal, cardinal, and decimal numbers in text streams.
bitmath converts and performs arithmetic on file sizes across SI and NIST prefix units (kB to YiB), with support for human-readable formatting, rich comparisons, and capacity math.
Converts ANSI escape sequences (color codes, formatting) from text to plain text output, removing all terminal styling.
However, verify the dual-license terms apply to your project, and consider whether you need active maintenance or bug fixes for your use case.
Parses and extracts data from HWP Document Format v5 files, with experimental support for conversion to OpenDocument (.odt) or plain text (.txt) formats.
However, be aware that the last release was in 2020, the experimental conversion features may not be production-ready, and AGPLv3+ licensing imposes copyleft…
ColorAide is a pure Python library for creating, converting, and manipulating colors across many color spaces with support for modern CSS color syntax and utilities like mixing, interpolation, and gamut mapping.
Adds an include directive to Python-Markdown that lets you embed the contents of other files into Markdown documents, with support for line-range selection and heading-depth adjustment.
Converts strings and other inputs to numbers (int, float, or real) with fast, flexible error handling and type-checking functions that outperform Python's built-in int() and float().
Converts plain-text punctuation to typographically correct HTML entities—replacing straight quotes with curly quotes, hyphens with em-dashes, and similar transformations for improved typography in web and document output.
Extracts text, tables, images, and metadata from 91+ file formats including PDFs, Office documents, and images, with native async/await support and multiple OCR backends.
Provides curated stop-word lists for multiple languages, enabling filtering of common words in natural language processing and text analysis tasks.
Install it if you need multilingual stop-word filtering for NLP or text preprocessing; the aging maintenance status is not a blocker for a stable, feature-complete…
Sumy extracts summaries from HTML pages or plain text using multiple automatic summarization algorithms (LSA, LexRank, Luhn, Edmundson) and provides evaluation tools for summary quality.
Panflute provides a Pythonic interface for writing Pandoc filters, allowing you to programmatically transform documents between markup formats by manipulating their abstract syntax trees.