crawlee
Crawlee for Python
Decision gist · record as of 2026-08-14
Yes. Crawlee is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It solves a real problem (reliable, parallel web scraping) with a modern async design and type hints. Choose it if you need both HTTP and browser-based crawling in one framework; if you only need simple HTTP requests, a lighter library may suffice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Playwright dependencies must be installed separately (playwright install) for browser-based crawling; BeautifulSoup extra required for HTML parsing crawler.
- Low friction installation with a pure-wheel distribution.
- Actively maintained with a release 8 days old and recent commits.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes, but there are no copyleft obligations.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 9,429 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 899,049 downloads/mo, #4,780 on PyPI
Alternatives
Verify before relying
pip install crawlee
import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main():
crawler = BeautifulSoupCrawler(max_requests_per_crawl=10)
@crawler.router.default_handler
async def handler(context: BeautifulSoupCrawlingContext):
await context.push_data({'url': context.request.url})
await crawler.run(['https://example.com'])
asyncio.run(main())- Whether impit (a runtime dependency) is a maintained HTTP client or has known issues.
- Performance characteristics under high concurrency or with large-scale crawls.
- Specific anti-bot evasion techniques and their effectiveness against modern protections.
What it is and what it does
Crawlee is a Python framework for building web scrapers and crawlers that abstracts away the complexity of HTTP requests, browser automation, parallel execution, and data persistence. It offers two main crawler types: BeautifulSoupCrawler for fast HTML parsing via HTTP, and PlaywrightCrawler for JavaScript-heavy sites that require a headless browser. Both crawlers handle retries, proxy rotation, session management, and request routing through a decorator-based handler system, storing results in a pluggable storage backend.
The library is built on asyncio for efficient concurrent crawling and includes type hints throughout for IDE support and static analysis. It targets developers who want reliable, maintainable scrapers without building infrastructure from scratch—handling bot detection evasion, URL queuing, error recovery, and data export out of the box. Core dependencies are minimal; optional extras (beautifulsoup, playwright) are installed only when needed.
Use it for
- Extract structured data from static HTML pages using BeautifulSoupCrawler for high-throughput scraping without browser overhead.
- Scrape JavaScript-rendered content by using PlaywrightCrawler to interact with dynamic pages and wait for content generation.
- Build a recursive web crawler that discovers and follows links across a domain with automatic URL queuing and deduplication.
- Rotate proxies and manage sessions automatically to avoid IP blocking and bot detection when scraping protected sites.
- Export crawled data to machine-readable formats (JSON, CSV) with persistent storage for fault-tolerant long-running jobs.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Crawlee is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It solves a real problem (reliable, parallel web scraping) with a modern async design and type hints. Choose it if you need both HTTP and browser-based crawling in one framework; if you only need simple HTTP requests, a lighter library may suffice.
Install
crawlee on PyPI
Before you install
Low friction installation with a pure-wheel distribution. Actively maintained with a release 8 days old and recent commits. Requires Python 3.10 or later. Optional extras (beautifulsoup, playwright) keep core dependencies lean; full-featured install via crawlee[all] adds browser automation capabilities.
Playwright dependencies must be installed separately (playwright install) for browser-based crawling; BeautifulSoup extra required for HTML parsing crawler.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes, but there are no copyleft obligations.
Quickstart
pip install crawlee
import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main():
crawler = BeautifulSoupCrawler(max_requests_per_crawl=10)
@crawler.router.default_handler
async def handler(context: BeautifulSoupCrawlingContext):
await context.push_data({'url': context.request.url})
await crawler.run(['https://example.com'])
asyncio.run(main())
Verify before relying
- Whether impit (a runtime dependency) is a maintained HTTP client or has known issues.
- Performance characteristics under high concurrency or with large-scale crawls.
- Specific anti-bot evasion techniques and their effectiveness against modern protections.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 12 packagesasync-timeoutcachetoolscoloramaimpitmore-itertoolsprotegopsutilpydantic-settingspydantictldextracttyping-extensionsyarl |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 899,049 / month, #4,780 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Libraries |
Evidence: crawlee-1.9.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “parallel web crawler”
- crawleeCrawlee is a web scraping and browser automation library that handles…
- icrawlericrawler is a lightweight, modular web crawler framework that…
- crawlerdetectIdentifies bots, crawlers, and spiders by analyzing user agents and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also Crawl4AI · icrawler · scrapling · Scrapy · zenrows · crawlerdetect · scrapy-playwright · LinkChecker · django-robots · pyppeteer