uroman
uroman is a universal romanizer. It converts text in any script to the standard Latin alphabet.
Decision gist · record as of 2026-08-14
Yes, if you need robust script-to-Latin conversion with language awareness. The low install friction, permissive license, and zero known vulnerabilities make it safe to adopt. However, the dormant maintenance status (no activity for 777 days) means you should verify that its romanization rules and Unicode support remain adequate for your use case and that you can maintain it yourself if needed.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Initial Uroman() constructor call takes about a second to load romanization data.
- Low install friction with a single runtime dependency (regex).
License · maintenance · safety
permissive license (permissive) — Permissive MIT-style license with an attribution requirement: publications using uroman must acknowledge its use and cite the original authors (Ulf Hermjakob, USC Information Sciences Institute, 2015-2020).
last release 2024-06-28 (777 days) · last repo commit 2024-07-26 · 249 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 659,136 downloads/mo, #5,464 on PyPI
Alternatives
Verify before relying
pip install uroman
import uroman as ur
uroman = ur.Uroman()
print(uroman.romanize_string('Игорь Стравинский'))
print(uroman.romanize_string('नेपाल', lcode='hin'))- Whether the dormant maintenance status (last commit 2024-07-26, no activity for 777 days) affects long-term compatibility with evolving Python or dependency ecosystems.
- Accuracy and completeness of romanization across all supported scripts and language codes beyond the documented examples.
What it is and what it does
uroman is a universal romanizer that converts text written in any script—Cyrillic, Greek, Arabic, Devanagari, Chinese, and many others—into standard Latin alphabet characters. It goes beyond simple character-by-character substitution by using context-aware m-to-n mappings and optional ISO-639-3 language codes to handle script-specific phonetic rules. For example, the same Cyrillic letter may romanize differently in Russian versus Ukrainian depending on the language code provided.
The package is designed to enable string-similarity comparisons across scripts without intermediate phonetic representations, and also converts numerals in various scripts to Western Arabic numerals. It offers both a command-line interface for batch processing and a Python API for programmatic use. The constructor loads romanization data once (taking about a second), after which multiple romanization calls are efficient. Output formats range from simple strings to detailed lattices with alternative romanizations and offset information.
Use it for
- Compare or deduplicate names and text across different writing systems in multilingual datasets without phonetic intermediate steps.
- Normalize user input from diverse scripts into Latin text for downstream NLP or machine translation pipelines.
- Convert digital numbers written in non-Latin scripts (e.g., Arabic, Chinese) to Western numerals for data processing.
- Build search indexes that match queries across scripts—e.g., finding 'Nepal' when the database contains नेपाल or نیپال.
- Preprocess multilingual corpora for linguistic research or computational linguistics tasks requiring consistent Latin representation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need robust script-to-Latin conversion with language awareness.
The low install friction, permissive license, and zero known vulnerabilities make it safe to adopt. However, the dormant maintenance status (no activity for 777 days) means you should verify that its romanization rules and Unicode support remain adequate for your use case and that you can maintain it yourself if needed.
Install
uroman on PyPI
Before you install
Low install friction with a single runtime dependency (regex). Maintenance is dormant—last release was 777 days ago and the repository shows no recent commits, though it remains unarchived and the package is actively downloaded.
Requires Python 3.10 or later. Initial Uroman() constructor call takes about a second to load romanization data.
License in practice
Permissive MIT-style license with an attribution requirement: publications using uroman must acknowledge its use and cite the original authors (Ulf Hermjakob, USC Information Sciences Institute, 2015-2020).
Quickstart
pip install uroman
import uroman as ur
uroman = ur.Uroman()
print(uroman.romanize_string('Игорь Стравинский'))
print(uroman.romanize_string('नेपाल', lcode='hin'))
Verify before relying
- Whether the dormant maintenance status (last commit 2024-07-26, no activity for 777 days) affects long-term compatibility with evolving Python or dependency ecosystems.
- Accuracy and completeness of romanization across all supported scripts and language codes beyond the documented examples.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packageregex |
| Maintenance | Dormant 777 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 659,136 / month, #5,464 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3 :: OnlyTopic :: Text ProcessingTopic :: Text Processing :: GeneralTopic :: Text Processing :: LinguisticTopic :: Utilities |
Evidence: uroman-1.3.1.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “romanize text any script”
- uromanConverts text in any script to Latin alphabet romanization,…
- hangul-romanizeConverts Korean Hangul text to romanized (Latin character)…
- pinyinConverts Chinese characters to pinyin (romanized Mandarin…
Give your agent the search over MCP, or paste the wish link into any chat.
More Utilities packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
See also hangul-romanize · transliterate · Unidecode · cyrtranslit · anyascii · cutlet · indic-numtowords · num2words · text-unidecode · number-parser