beautifulsoup4
Screen-scraping library
Decision gist · record as of 2026-08-14
Yes. Beautiful Soup is a mature, widely-used standard for web scraping and HTML/XML parsing in Python. It has low install friction, active maintenance, a permissive MIT license, no known vulnerabilities, and ranks in the top 100 most-downloaded PyPI packages. Install it if you need to parse or extract data from markup documents.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.7.0 or later; relies on an underlying HTML/XML parser.
- Low friction install with only two lightweight runtime dependencies.
- Active maintenance with a recent release 68 days ago.
License · maintenance · safety
MIT License (permissive) — MIT License permits commercial and private use with minimal restrictions—you may use, modify, and distribute freely provided you retain the license notice.
last release 2026-06-07 (68 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 432,133,462 downloads/mo, #96 on PyPI
Alternatives
Verify before relying
pip install beautifulsoup4
from beautifulsoup4 import BeautifulSoup
soup = BeautifulSoup("<html><body>content</body></html>")
print(soup.find(string="content"))- Whether the package includes or recommends additional parser libraries for improved performance or robustness.
- Performance characteristics when parsing very large documents or handling deeply nested structures.
- Specific import path and module structure for the package API.
What it is and what it does
Beautiful Soup is a screen-scraping library that takes raw HTML or XML markup and builds a navigable parse tree you can query and modify using straightforward Python idioms. It sits on top of a parser and abstracts away the low-level parsing details, letting you focus on finding and extracting the data you need.
The library is designed for web scraping and data extraction tasks where you need to pull structured information from markup documents. It handles malformed HTML gracefully, supports both HTML and XML modes, and provides methods to locate elements by tag, attribute, or content. It depends on soupsieve and typing-extensions at runtime.
Use it for
- Extract product listings, prices, or metadata from e-commerce websites for price comparison or market research.
- Scrape news articles, headlines, or publication metadata from news sites for aggregation or analysis.
- Parse HTML documentation or API responses to automate data collection from web-based sources.
- Clean and normalize malformed HTML from legacy systems or user-generated content before processing.
- Build web crawlers that need to navigate and extract links or structured data from multiple pages.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Beautiful Soup is a mature, widely-used standard for web scraping and HTML/XML parsing in Python. It has low install friction, active maintenance, a permissive MIT license, no known vulnerabilities, and ranks in the top 100 most-downloaded PyPI packages. Install it if you need to parse or extract data from markup documents.
Install
beautifulsoup4 on PyPI
Before you install
Low friction install with only two lightweight runtime dependencies. Active maintenance with a recent release 68 days ago.
Requires Python 3.7.0 or later; relies on an underlying HTML/XML parser.
License in practice
MIT License permits commercial and private use with minimal restrictions—you may use, modify, and distribute freely provided you retain the license notice.
Quickstart
pip install beautifulsoup4
from beautifulsoup4 import BeautifulSoup
soup = BeautifulSoup("<html><body>content</body></html>")
print(soup.find(string="content"))
Verify before relying
- Whether the package includes or recommends additional parser libraries for improved performance or robustness.
- Performance characteristics when parsing very large documents or handling deeply nested structures.
- Specific import path and module structure for the package API.
Package facts
| License | MIT License permissive |
| Python support | Supports the current Python release >=3.7.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagessoupsievetyping-extensions |
| Maintenance | Actively maintained 68 days since the last release |
| First released | |
| Downloads | 432,133,462 / month, #96 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Topic :: Software Development :: Libraries :: Python ModulesTopic :: Text Processing :: Markup :: HTMLTopic :: Text Processing :: Markup :: SGMLTopic :: Text Processing :: Markup :: XML |
Evidence: beautifulsoup4-4.15.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also BeautifulSoup · bs4 · soup2dict · prettierfier · soupsieve · googlesearch-python · turbohtml · feedparser-sgmllib · webargs · readabilipy