$npx skillfedfor your agent

PyArabic

Arabic text tools for Python

With conditionsPyPI LinguisticReleased Jun 2022152.3K downloads / moGPLPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — PyArabic-0.6.15-py3-none-any.whl
v0.6.15 · released 2022-06-18 · 1 runtime deps: six

Yes, if you need stable Arabic text manipulation and the existing feature set meets your requirements. The low install friction and lack of security vulnerabilities make it safe to use. However, the abandoned maintenance status means no bug fixes or updates will be forthcoming—do not install if you expect ongoing support or compatibility with future Python versions. Suitable for production use only in stable, unchanging workflows.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python with Unicode support (native in Python 3; Python 2 requires proper encoding declarations).
  • Arabic text must be defined with u'' prefix or UTF-8 encoding declaration.
  • Install friction is low; the package depends only on six.

License · maintenance · safety

GPL (copyleft) — PyArabic is licensed under GPL (copyleft). Any derivative work or distribution must also be released under GPL; proprietary or closed-source projects cannot incorporate it without legal risk.

last release 2022-06-18 (1518 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 152,348 downloads/mo, #10,905 on PyPI

Verify before relying

pip install pyarabic

import pyarabic.araby as araby
text = u'السلام عليكم'
stripped = araby.strip_tashkeel(text)
  • Whether the package works reliably with modern Python versions (3.9+) given its abandoned status.
  • Performance characteristics when processing large Arabic corpora.
  • Compatibility with recent versions of the six dependency.
Same gist for agents: .md · .json

What it is and what it does

PyArabic is a specialized library for processing Arabic text in Python. It provides a collection of functions organized into modules—araby.py for general text operations (stripping diacritics, tokenization, character classification), number.py for converting between numerals and Arabic words, and named.py for named-entity recognition. The library works with Unicode-encoded Arabic strings and handles Arabic-specific challenges like ligatures, hamza variants, and diacritical marks (harakat).

The package is designed for developers building Arabic natural language processing pipelines, text normalization workflows, or linguistic analysis tools. It depends only on six and installs with low friction. However, the project is abandoned—the last release was over four years ago—so it receives no maintenance, bug fixes, or updates. It is suitable only for stable use cases where the existing feature set is sufficient.

Use it for

  • Strip diacritical marks from Arabic text for stemming or lemmatization in search or NLP pipelines.
  • Tokenize Arabic documents into words or sentences for text analysis or corpus processing.
  • Normalize Arabic script variants (hamza, ligatures) to standardize text before comparison or indexing.
  • Convert numeric values to Arabic words or extract numeric phrases from Arabic text for document parsing.
  • Classify and detect Arabic letters and character groups for text validation or linguistic analysis.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need stable Arabic text manipulation and the existing feature set meets your requirements.

The low install friction and lack of security vulnerabilities make it safe to use. However, the abandoned maintenance status means no bug fixes or updates will be forthcoming—do not install if you expect ongoing support or compatibility with future Python versions. Suitable for production use only in stable, unchanging workflows.

Install

pyarabic on PyPI

Before you install

Install friction is low; the package depends only on six. However, maintenance is abandoned—the last release was 1518 days ago, and no recent commits or updates are evident. Use only if the existing functionality meets your needs without expecting bug fixes or feature additions.

Requires Python with Unicode support (native in Python 3; Python 2 requires proper encoding declarations). Arabic text must be defined with u'' prefix or UTF-8 encoding declaration.

License in practice

PyArabic is licensed under GPL (copyleft). Any derivative work or distribution must also be released under GPL; proprietary or closed-source projects cannot incorporate it without legal risk.

Quickstart

pip install pyarabic

import pyarabic.araby as araby
text = u'السلام عليكم'
stripped = araby.strip_tashkeel(text)

Verify before relying

  • Whether the package works reliably with modern Python versions (3.9+) given its abandoned status.
  • Performance characteristics when processing large Arabic corpora.
  • Compatibility with recent versions of the six dependency.

Package facts

LicenseGPL copyleft
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
six
MaintenanceAbandoned 1,518 days since the last release
First released
Downloads152,348 / month, #10,905 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersNatural Language :: ArabicOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 3Topic :: Text Processing :: Linguistic

Evidence: PyArabic-0.6.15-py3-none-any.whl

Tags

Capabilities
arabic text processingarabic letter classificationremove arabic diacriticsarabic tokenizationarabic string normalizationarabic nlp utilitiesarabic character detection
Topics
arabic-nlptext-processingabandoned-but-stable

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “arabic text processing”

  • PyArabicPyArabic provides functions to manipulate Arabic text and…
  • arabic-reshaperReshapes Arabic text characters into their correct contextual forms…
  • python-bidiConverts bidirectional text (mixed left-to-right and right-to-left…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also arabic-reshaper · normality · indic-nlp-library · fold-to-ascii · jieba3k · jieba · confusables · segments · w3lib · zalgolib