scrapegraph-py
Official Python SDK for ScrapeGraph AI API
Decision gist · record as of 2026-08-14
Yes, if you need web scraping with minimal infrastructure setup and can absorb the API costs. The SDK is actively maintained, has low install friction, and offers a straightforward path to production. No if you require on-premises data handling, want to avoid per-request billing, or need fine-grained control over LLM selection and browser configuration—use the open-source scrapegraphai library instead.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires SGAI_API_KEY environment variable or explicit api_key parameter; the API is a paid cloud service, not free.
- Low install friction with only two runtime dependencies (httpx, pydantic).
- Active maintenance with a release 14 days ago and recent commits.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial use, modification, and distribution with minimal restrictions. The SDK itself is open-source, though the underlying API service is a paid cloud offering.
last release 2026-07-31 (14 days) · last repo commit 2026-07-31 · 83 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 294,833 downloads/mo, #7,939 on PyPI
Alternatives
Verify before relying
pip install scrapegraph-py
from scrapegraph_py import ScrapeGraphAI
sgai = ScrapeGraphAI() # reads SGAI_API_KEY from env
result = sgai.scrape("https://example.com")
if result.status == "success":
print(result.data["results"]["markdown"]["data"])- Rate limits, quota tiers, or per-request costs for the managed API service.
- Whether async client supports all methods mentioned in sync documentation.
- Supported Python versions beyond 3.12 and 3.13 (requires_python states >=3.12 but classifiers list only 3.12 and 3.13).
What it is and what it does
scrapegraph-py is an official Python client for ScrapeGraphAI's managed web scraping API. It wraps a cloud service that handles web page fetching, JavaScript rendering, LLM-based content extraction, and anti-bot measures on your behalf. You send a URL and optional format/extraction prompt; the API returns structured results (markdown, HTML, JSON, screenshots, summaries, or extracted data) without managing browsers, proxies, or LLM keys yourself.
The SDK provides both sync and async interfaces for six main operations: scrape (fetch pages in multiple formats), extract (AI-powered structured data from URLs or raw HTML), search (web search with optional extraction), crawl (multi-page site traversal with depth/link limits), monitor (scheduled change detection via cron), and history (request audit trail). All methods return an ApiResult wrapper with status, data, error, and elapsed_ms fields—no exceptions to catch. It differs from the open-source scrapegraphai library, which runs locally and requires you to manage LLMs, browsers, and infrastructure.
Use it for
- Scrape product listings and prices from e-commerce sites without writing site-specific parsers.
- Extract structured data (JSON schema) from unstructured web content using AI prompts.
- Monitor competitor websites or pricing pages on a schedule and receive webhook notifications.
- Crawl documentation or blog sites to collect all pages in markdown format for indexing.
- Search the web and extract key information from results in a single API call.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need web scraping with minimal infrastructure setup and can absorb the API costs.
The SDK is actively maintained, has low install friction, and offers a straightforward path to production. No if you require on-premises data handling, want to avoid per-request billing, or need fine-grained control over LLM selection and browser configuration—use the open-source scrapegraphai library instead.
Install
scrapegraph-py on PyPI
Before you install
Low install friction with only two runtime dependencies (httpx, pydantic). Active maintenance with a release 14 days ago and recent commits. Beta status and early project age (first release November 2024) mean the API surface may still shift.
Requires SGAI_API_KEY environment variable or explicit api_key parameter; the API is a paid cloud service, not free.
License in practice
MIT license permits commercial use, modification, and distribution with minimal restrictions. The SDK itself is open-source, though the underlying API service is a paid cloud offering.
Quickstart
pip install scrapegraph-py
from scrapegraph_py import ScrapeGraphAI
sgai = ScrapeGraphAI() # reads SGAI_API_KEY from env
result = sgai.scrape("https://example.com")
if result.status == "success":
print(result.data["results"]["markdown"]["data"])
Verify before relying
- Rate limits, quota tiers, or per-request costs for the managed API service.
- Whether async client supports all methods mentioned in sync documentation.
- Supported Python versions beyond 3.12 and 3.13 (requires_python states >=3.12 but classifiers list only 3.12 and 3.13).
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.12 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packageshttpxpydantic |
| Maintenance | Actively maintained 14 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 294,833 / month, #7,939 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Internet :: WWW/HTTP :: Indexing/SearchTopic :: Software Development :: Libraries :: Python ModulesTyping :: Typed |
Evidence: scrapegraph_py-2.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ai-powered web scraper”
- scrapegraph-pyClient SDK for the ScrapeGraphAI managed API, enabling web scraping,…
- scrapfly-sdkPython SDK for the Scrapfly web scraping service, providing access to…
- goose3Extracts article text, metadata, images, and embedded videos from web…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also brightdata-sdk · spider-client · firecrawl · firecrawl-py · Crawl4AI · scrapfly-sdk · scrapingbee · tavily-python · googlesearch-python · tavily-cli