RTFDE
A library for extracting HTML content from RTF encapsulated HTML as commonly found in the exchange MSG email format.
Decision gist · record as of 2026-08-14
Yes, if you need to extract HTML or text from RTF-encapsulated .msg email bodies. The library is narrow in scope, low friction to install, and has no known vulnerabilities. The aging maintenance status and small community (8 stars) mean it is unlikely to gain new features, but the core task is stable and unlikely to require frequent updates. Suitable for production use in email processing pipelines where RTF de-encapsulation is a known requirement.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >= 3.8.
- Input must be raw RTF bytes (typically from a .msg file body).
- Low friction: pure Python wheel with only two runtime dependencies (lark and oletools).
License · maintenance · safety
copyleft license (copyleft) — Licensed under LGPLv3 (copyleft). You may use and modify the library freely, but any derivative work must also be distributed under LGPLv3; static linking or bundling into proprietary software requires careful legal review.
last release 2025-12-09 (248 days) · last repo commit 2025-12-09 · 8 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 6,577,070 downloads/mo, #1,886 on PyPI
Alternatives
Verify before relying
from RTFDE.deencapsulate import DeEncapsulator
with open('rtf_file', 'rb') as fp:
raw_rtf = fp.read()
rtf_obj = DeEncapsulator(raw_rtf)
rtf_obj.deencapsulate()
if rtf_obj.content_type == 'html':
print(rtf_obj.html)
else:
print(rtf_obj.text)- How well does de-encapsulation handle modern Exchange RTF variants and edge cases beyond the documented known issues?
- Performance characteristics with large RTF payloads or batch processing of many .msg files.
- Compatibility with oletools versions and whether oletools updates could affect behavior.
What it is and what it does
RTFDE is a Python library that reverses Microsoft Exchange's RTF encapsulation of email body content. When Outlook or Exchange stores HTML or plain text email bodies, they wrap them in RTF format; this library extracts the original HTML or text from that wrapper. It depends on lark (for parsing) and oletools (for RTF/OLE handling) and is designed specifically for .msg file processing.
The library handles two main tasks: extracting HTML from RTF-encapsulated HTML, and extracting plain text from RTF-encapsulated text. It fully unquotes the extracted content, which means escaped sequences (like Quoted-Printable encoding) are decoded. Known limitations include inability to integrate attachments back into HTML bodies and no support for extracting plain text from RTF-encapsulated HTML (you would use another HTML parser for that).
Use it for
- Parse .msg email files and recover the original HTML body content instead of the RTF wrapper.
- Batch-process Exchange-exported emails to extract and re-render body content as clean HTML or text.
- Build email archival or migration tools that need to normalize RTF-wrapped content from older Outlook exports.
- Integrate with oletools-based email forensics workflows to de-encapsulate message bodies.
- Convert RTF-encapsulated plain text from .msg files back to readable text for indexing or search.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to extract HTML or text from RTF-encapsulated .msg email bodies.
The library is narrow in scope, low friction to install, and has no known vulnerabilities. The aging maintenance status and small community (8 stars) mean it is unlikely to gain new features, but the core task is stable and unlikely to require frequent updates. Suitable for production use in email processing pipelines where RTF de-encapsulation is a known requirement.
Install
rtfde on PyPI
Before you install
Low friction: pure Python wheel with only two runtime dependencies (lark and oletools). Maintenance is aging—last release was 248 days ago and the repository has minimal activity (8 stars), but it remains archived=false and the codebase is stable for its narrow scope.
Requires Python >= 3.8. Input must be raw RTF bytes (typically from a .msg file body).
License in practice
Licensed under LGPLv3 (copyleft). You may use and modify the library freely, but any derivative work must also be distributed under LGPLv3; static linking or bundling into proprietary software requires careful legal review.
Quickstart
from RTFDE.deencapsulate import DeEncapsulator
with open('rtf_file', 'rb') as fp:
raw_rtf = fp.read()
rtf_obj = DeEncapsulator(raw_rtf)
rtf_obj.deencapsulate()
if rtf_obj.content_type == 'html':
print(rtf_obj.html)
else:
print(rtf_obj.text)
Verify before relying
- How well does de-encapsulation handle modern Exchange RTF variants and edge cases beyond the documented known issues?
- Performance characteristics with large RTF payloads or batch processing of many .msg files.
- Compatibility with oletools versions and whether oletools updates could affect behavior.
Package facts
| License | copyleft license copyleft |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packageslarkoletools |
| Maintenance | Aging 248 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 6,577,070 / month, #1,886 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersLicense :: OSI Approved :: GNU Lesser General Public License v3 (LGPLv3)Operating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Communications :: Email :: FiltersTopic :: Text Processing :: FiltersTopic :: Text Processing :: Markup :: HTML |
Evidence: rtfde-0.1.2.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “rtf de-encapsulation”
- RTFDEExtracts encapsulated HTML and plain text content from RTF bodies in…
- rtfunicodeEncodes Unicode strings to RTF 1.5 command sequences, registering a…
- striprtfConverts Rich Text Format (RTF) files to plain text, stripping…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also rtfparse · tnefparse · extract-msg · striprtf · msg-parser · quotequail · mail-parser · python-oxmsg · pyth3 · markdown2