mwxml
A set of utilities for processing MediaWiki XML dump data.
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has no known vulnerabilities, installs with low friction, and solves a real problem for anyone working with MediaWiki dumps. The MIT license imposes no restrictions. Install it if you need to process wiki database exports; skip it if you're not working with MediaWiki XML data.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with a pure-Python wheel distribution.
- Active maintenance with a recent release (128 days ago) and ongoing repository activity.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal legal constraints.
last release 2026-04-08 (128 days) · last repo commit 2026-04-09 · 63 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 249,574 downloads/mo, #8,648 on PyPI
Alternatives
Verify before relying
pip install mwxml
import mwxml
dump = mwxml.Dump.from_file(open("dump.xml"))
for page in dump:
for revision in page:
print(revision.id)- Whether the distributed processing strategy (map function) requires additional setup or dependencies beyond the listed runtime packages.
- Performance characteristics and memory overhead when processing very large dump files.
- Python version compatibility details, as requires_python is unspecified despite Python 3 Only classifier.
What it is and what it does
mwxml is a Python library for processing MediaWiki XML database dumps—the raw XML exports that contain Wikipedia and other wiki content. It addresses two core problems: the complexity of parsing large XML files and the performance overhead of naive approaches. The library provides a simple iterator interface that streams through dump files without loading them entirely into memory, making it practical to process dumps that would otherwise exhaust available RAM.
The package also supports distributed processing, allowing you to work with multiple dump files in parallel. It depends on mwtypes, mwcli, para, and jsonschema to handle MediaWiki-specific data types, command-line utilities, parallelization, and schema validation. It's designed for developers and researchers who need to extract, analyze, or transform content from wiki database exports.
Use it for
- Extract revision history and metadata from Wikipedia or other MediaWiki dumps for linguistic or historical analysis.
- Process multiple large XML dump files in parallel to build indexes or data warehouses from wiki content.
- Stream through a dump file to filter and extract specific pages or revisions without loading the entire file into memory.
- Analyze edit patterns, contributor activity, or content evolution by iterating over revisions in a structured way.
- Build data pipelines that consume wiki dumps as input for downstream NLP or machine learning workflows.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has no known vulnerabilities, installs with low friction, and solves a real problem for anyone working with MediaWiki dumps. The MIT license imposes no restrictions. Install it if you need to process wiki database exports; skip it if you're not working with MediaWiki XML data.
Install
mwxml on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release (128 days ago) and ongoing repository activity.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal legal constraints.
Quickstart
pip install mwxml
import mwxml
dump = mwxml.Dump.from_file(open("dump.xml"))
for page in dump:
for revision in page:
print(revision.id)
Verify before relying
- Whether the distributed processing strategy (map function) requires additional setup or dependencies beyond the listed runtime packages.
- Performance characteristics and memory overhead when processing very large dump files.
- Python version compatibility details, as requires_python is unspecified despite Python 3 Only classifier.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesmwtypesmwcliparajsonschema |
| Maintenance | Actively maintained 128 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 249,574 / month, #8,648 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Environment :: Other EnvironmentIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyTopic :: Scientific/EngineeringTopic :: Software Development :: Libraries :: Python ModulesTopic :: Text Processing :: GeneralTopic :: Text Processing :: LinguisticTopic :: Utilities |
Evidence: mwxml-0.3.8-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “mediawiki xml parsing”
- mwxmlEfficiently parse and stream-process MediaWiki XML database dumps…
- mwtypesProvides standardized Python classes for MediaWiki data types,…
- mwcliProvides helper functions and classes for building command-line…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also mwtypes · mwcli · mwclient · mwparserfromhell · wikipedia · wikitextparser · crick · Wikipedia-API · m3u8 · xmljson