Subcategories
Packages
Computes the edit distance (Levenshtein distance) between two sequences using a fast C++ and Cython implementation, supporting strings and any hashable iterables.
However, be aware the repository is archived; no new features or security updates will be released.
Parsy is a parser combinator library that lets you build complex text parsers by combining small, reusable parser functions together in a declarative style.
Install it if you need to parse structured text and prefer a functional, composable approach over regex or hand-rolled parsing logic.
A Python client for the Chunkr document intelligence API that uploads files and images for processing, supporting both synchronous and asynchronous workflows.
However, you depend entirely on the Chunkr backend's availability and the package's aging maintenance status suggests reduced active development—verify that the…
jaconv converts between Japanese character types: Hiragana, Katakana, half-width (Hankaku) and full-width (Zenkaku) characters, plus transliteration to and from romanized alphabet.
Install it if you work with Japanese text preprocessing, normalization, or transliteration.
Converts marked-up text between formats including plain text, XHTML, RTF, and PDF, preserving basic structure like paragraphs, headings, lists, and simple tables.
A fast text parser library that lets you define grammars using token specs and parsing rules, then parse text into structured parse trees.
Install it if you're building a parser and speed is a priority; skip it if you need error recovery, detailed diagnostics, or support for older Python versions.
Parses human-readable time expressions (like '2h32m', '1.2 minutes', '2 days, 4:13:02') into seconds or timedelta objects.
Install it if you need to accept time expressions from users or configuration; skip it if you only work with numeric seconds or strict time formats.
Generates Word documents (.docx) from templates by embedding Jinja2 tags into a Word document and populating them with context variables at runtime.
Install it if you need to generate Word documents from templates with dynamic content.
pypandoc_binary wraps pandoc, a universal document converter, and bundles the pandoc binary so you can convert between document formats without a separate pandoc installation.
Provides type hints for the Pygments syntax highlighting library, enabling static type checkers to validate code that uses Pygments.
rjsmin minifies JavaScript code by removing unnecessary characters while preserving functionality, implemented in Python with optional C acceleration for runtime use.
However, consider the 306-day release gap when evaluating long-term support; if you require active maintenance or are starting a new project, evaluate whether a more…
Applies ANSI true color and text styles (bold, italic, underline, etc.) to terminal output using hex color codes or named colors.
Install it if you need to add true color and text effects to terminal output; skip it if your application doesn't interact with terminals or if you're already using a…
RCSSmin minifies CSS by removing spaces, comments, and unnecessary characters while preserving CSS semantics and supporting common CSS hacks.
However, the 306-day gap since the last release and aging maintenance status suggest limited active development; verify that the package still meets your platform and…
Converts strings to snake_case by normalizing case, stripping punctuation, and handling camelCase or PascalCase inputs.
However, do not use it if you require ongoing maintenance, support for recent Python versions, or security updates—consider a maintained alternative instead.
ast-grep-cli is a command-line tool for searching, linting, and rewriting code using abstract syntax tree patterns instead of text matching.
Parses and applies patch files in multiple diff formats (unified, context, normal, ed, rcs) and from several source control systems (Git, SVN, CVS).
Extracts or replaces keywords in text using the FlashText algorithm, which is based on Aho-Corasick and Trie data structures and designed to be faster than regex for bulk keyword operations.
pdfrw reads, writes, and manipulates PDF files in pure Python, supporting operations like merging, subsetting, rotating, and metadata modification without external dependencies.
Detects and replaces personally identifiable information (names, emails, phone numbers, credit cards, dates of birth, social security numbers, and more) in free text with anonymized placeholders.
However, maintenance is dormant, so if you need active bug fixes or feature development, evaluate whether the package's current detector coverage meets your needs.
Provides utility functions for common text operations including wrapping, substitution, trimming, stripping, prefix/suffix removal, indentation, and case-insensitive comparison.
Install it if you need a collection of common text-processing helpers and are comfortable with a Python 3.10+ requirement.
Translates text between languages using multiple free translation services (Google, Microsoft, DeepL, Yandex, and others) with automatic language detection and batch processing support.
However, do not rely on it for production systems where translation service availability is critical—the dormant status means breakage from service changes will not…
Jieba segments Chinese text into words using multiple algorithms (precise, full, and search-engine modes) and supports both simplified and traditional Chinese with custom dictionary injection.
However, install it only if you are working with Chinese text and can verify compatibility with your Python version—the package is dormant and may not work on very…
Casefy converts strings between different casing conventions (camelCase, snake_case, kebab-case, PascalCase, CONST_CASE, and others) with Unicode support and no external dependencies.
Parses Lucene Query DSL syntax into an abstract syntax tree, enabling inspection, analysis, and transformation of search queries for use with Elasticsearch or custom backends.
Install it if you need to accept Lucene syntax from users and convert it to Elasticsearch or apply custom query logic.
Converts terminal output with ANSI color codes to HTML or LaTeX, preserving colors and formatting for display in browsers or documents.
Install it if you need to convert ANSI-colored terminal output to HTML or LaTeX.
ProperDocs is a static site generator that converts Markdown documentation into HTML, configured through a single YAML file and extensible with themes and plugins.
Install it if you need a straightforward static documentation generator; verify the plugin ecosystem and programmatic API match your workflow before committing to it…
Cleans and normalizes US addresses to USPS and RESO standards, converting them to a consistent dictionary format with uppercase fields and standardized abbreviations.
However, verify the license status before production use, as it is currently marked unclear in the metadata.
Jsonnet is a data templating language that evaluates Jsonnet code to produce JSON output, with Python bindings to the original C++ implementation.
However, do not use it to evaluate untrusted Jsonnet code without external sandboxing—the C++ implementation is not hardened.
Parses HogQL query expressions and Hog programs into abstract syntax trees, available as a native C++ extension for Python on macOS and Linux.
Parses HogQL query expressions into JSON AST format via a hand-rolled Rust parser, delivering 15–50× performance gains over the C++ ANTLR reference implementation.
Provides a collection of Jinja2 template filters ported from Ansible, including encoding, hashing, path manipulation, JSON/YAML conversion, and regex operations, without requiring Ansible core.
However, do not install if your project is proprietary or closed-source (GPL3 is a blocker), or if you require active maintenance and compatibility updates—the…
A Sphinx extension that adds tabbed content blocks to HTML documentation, supporting basic tabs, grouped tabs with synchronized selection, and syntax-highlighted code tabs.
Cheetah3 is a template engine and code generation tool that processes templates to produce output in various languages including Python, C++, Java, and SQL.
BM25S implements the BM25 ranking algorithm in pure Python with Numpy, enabling fast document retrieval and ranking based on text queries.
Install it if you need to rank documents by text relevance in Python without external services.
Computes the Levenshtein distance between two strings—the minimum number of single-character edits needed to transform one string into another.
Converts Chinese characters to pinyin (romanized pronunciation) with support for multiple styles, heteronyms, and tone marks.
Install it if you need to work with Chinese character romanization; the only caveat is that accuracy depends on proper word segmentation and may require custom…
Parses and serializes iCalendar and vCard files, converting them to and from Python data structures while handling relevant encodings.
flpc wraps the Rust regex crate to provide faster regular expression matching, searching, and substitution operations than Python's native `re` module, with a similar but slightly modified API.
Install only if you have a short-term, isolated use case where you can afford to fork or maintain it yourself, and you've verified that the performance gain justifies…
Provides a familiar JSON Lines (ndjson) parser and writer with an API matching Python's built-in json and pickle modules, supporting streaming and file I/O for newline-delimited JSON data.
Detects and censors profanity in text, including leetspeak variations like p0rn and h4NDjob, using fast string comparison instead of regex.
However, maintenance is dormant since 2020-11-02, so the wordlist and codebase are not actively updated.