$npx skillfedfor your agent

crawlee

Crawlee for Python

Worth itPyPI LibrariesReleased Aug 2026899.0K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — crawlee-1.9.1-py3-none-any.whl
v1.9.1 · released 2026-08-06 · Python >=3.10 · 12 runtime deps: async-timeout, cachetools, colorama, impit, more-itertools, protego, psutil, pydantic-settings

Yes. Crawlee is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It solves a real problem (reliable, parallel web scraping) with a modern async design and type hints. Choose it if you need both HTTP and browser-based crawling in one framework; if you only need simple HTTP requests, a lighter library may suffice.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Playwright dependencies must be installed separately (playwright install) for browser-based crawling; BeautifulSoup extra required for HTML parsing crawler.
  • Low friction installation with a pure-wheel distribution.
  • Actively maintained with a release 8 days old and recent commits.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes, but there are no copyleft obligations.

last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 9,429 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 899,049 downloads/mo, #4,780 on PyPI

Verify before relying

pip install crawlee

import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext

async def main():
    crawler = BeautifulSoupCrawler(max_requests_per_crawl=10)
    @crawler.router.default_handler
    async def handler(context: BeautifulSoupCrawlingContext):
        await context.push_data({'url': context.request.url})
    await crawler.run(['https://example.com'])

asyncio.run(main())
  • Whether impit (a runtime dependency) is a maintained HTTP client or has known issues.
  • Performance characteristics under high concurrency or with large-scale crawls.
  • Specific anti-bot evasion techniques and their effectiveness against modern protections.
Same gist for agents: .md · .json

What it is and what it does

Crawlee is a Python framework for building web scrapers and crawlers that abstracts away the complexity of HTTP requests, browser automation, parallel execution, and data persistence. It offers two main crawler types: BeautifulSoupCrawler for fast HTML parsing via HTTP, and PlaywrightCrawler for JavaScript-heavy sites that require a headless browser. Both crawlers handle retries, proxy rotation, session management, and request routing through a decorator-based handler system, storing results in a pluggable storage backend.

The library is built on asyncio for efficient concurrent crawling and includes type hints throughout for IDE support and static analysis. It targets developers who want reliable, maintainable scrapers without building infrastructure from scratch—handling bot detection evasion, URL queuing, error recovery, and data export out of the box. Core dependencies are minimal; optional extras (beautifulsoup, playwright) are installed only when needed.

Use it for

  • Extract structured data from static HTML pages using BeautifulSoupCrawler for high-throughput scraping without browser overhead.
  • Scrape JavaScript-rendered content by using PlaywrightCrawler to interact with dynamic pages and wait for content generation.
  • Build a recursive web crawler that discovers and follows links across a domain with automatic URL queuing and deduplication.
  • Rotate proxies and manage sessions automatically to avoid IP blocking and bot detection when scraping protected sites.
  • Export crawled data to machine-readable formats (JSON, CSV) with persistent storage for fault-tolerant long-running jobs.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Crawlee is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It solves a real problem (reliable, parallel web scraping) with a modern async design and type hints. Choose it if you need both HTTP and browser-based crawling in one framework; if you only need simple HTTP requests, a lighter library may suffice.

Install

crawlee on PyPI

Before you install

Low friction installation with a pure-wheel distribution. Actively maintained with a release 8 days old and recent commits. Requires Python 3.10 or later. Optional extras (beautifulsoup, playwright) keep core dependencies lean; full-featured install via crawlee[all] adds browser automation capabilities.

Playwright dependencies must be installed separately (playwright install) for browser-based crawling; BeautifulSoup extra required for HTML parsing crawler.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes, but there are no copyleft obligations.

Quickstart

pip install crawlee

import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext

async def main():
    crawler = BeautifulSoupCrawler(max_requests_per_crawl=10)
    @crawler.router.default_handler
    async def handler(context: BeautifulSoupCrawlingContext):
        await context.push_data({'url': context.request.url})
    await crawler.run(['https://example.com'])

asyncio.run(main())

Verify before relying

  • Whether impit (a runtime dependency) is a maintained HTTP client or has known issues.
  • Performance characteristics under high concurrency or with large-scale crawls.
  • Specific anti-bot evasion techniques and their effectiveness against modern protections.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
12 packages
async-timeoutcachetoolscoloramaimpitmore-itertoolsprotegopsutilpydantic-settingspydantictldextracttyping-extensionsyarl
MaintenanceActively maintained 8 days since the last release
Last repo commit
First released
Downloads899,049 / month, #4,780 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Libraries

Evidence: crawlee-1.9.1-py3-none-any.whl

Tags

Capabilities
web scraping library pythonheadless browser automationparallel web crawlerhttp scraper with javascript supportasync web scraping frameworkproxy rotation crawlerurl queue management
Topics
web-scrapingasync-crawlerbrowser-automation
PyPI keywords
apifyautomationchromecrawleecrawlerheadlessscraperscraping

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “parallel web crawler”

  • crawleeCrawlee is a web scraping and browser automation library that handles…
  • icrawlericrawler is a lightweight, modular web crawler framework that…
  • crawlerdetectIdentifies bots, crawlers, and spiders by analyzing user agents and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also Crawl4AI · icrawler · scrapling · Scrapy · zenrows · crawlerdetect · scrapy-playwright · LinkChecker · django-robots · pyppeteer