Packages
Provides stemming algorithms for 34 languages, reducing word variants to a common stem for text search and indexing applications.
Install it if you need stemming for search or text indexing.
Represents and manipulates IPv4, IPv6, MAC addresses and related network objects; supports CIDR notation, subnetting, set operations, IANA lookups, and DNS reverse generation.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
Aggregates web search results from multiple search engines (Bing, DuckDuckGo, Google, Brave, and others) through a single Python API, with optional distributed caching via DHT network.
Install it if you need programmatic multi-engine search; skip it only if you require a single specific search engine's API or have strict Windows-only requirements.
Retrieves search results from DuckDuckGo for text, images, videos, and news without requiring an API key.
Install only if you have an existing dependency or need to maintain legacy code.
A Python HTTP client for interacting with Algolia's search API, supporting both synchronous and asynchronous operations with built-in request handling and response parsing.
PyStemmer provides word stemming algorithms for multiple languages, reducing words to their morphological base form to improve search and information retrieval.
A Python SDK for web scraping, crawling, searching, and extracting structured data from websites and research papers via the Firecrawl API, returning results as clean Markdown, HTML, or typed objects.
Official Python client for SerpApi that queries multiple search engines and data sources (Google, Bing, Baidu, Yahoo, eBay, YouTube, and others) and returns structured results as dictionary-like objects.
Install it if you need programmatic access to search results or structured data from multiple sources.
A Python wrapper for Censys APIs that lets you search internet-wide data on hosts, certificates, and services, manage attack surface assets, and run queries from the command line.
However, review the deprecation notice carefully: if you are starting a new project, evaluate censys-platform as your long-term choice.
Python SDK for the Spider Cloud API that enables website scraping, crawling, link extraction, screenshot capture, and LLM-compatible data collection through a managed cloud service.
However, it requires a Spider Cloud account and API key—this is a client library for a paid service, not a standalone scraper.
Automate programmatic interaction with HTTP web servers by simulating a stateful browser—fill forms, follow links, manage history, and parse HTML without a GUI.
However, consider that maintenance is aging—last release was 840 days ago—so evaluate whether the package meets your Python version and modern web compatibility needs…
pysolr is a lightweight Python client for Apache Solr that enables querying, indexing, and deleting documents in a Solr search server, with support for advanced features like highlighting, faceting, and SolrCloud.
Web Forager is a search-and-fetch toolkit for AI agents that retrieves web content via DuckDuckGo, news search, and Jina Reader, then converts results into LLM-friendly formats.
The main gotcha is Python version support (3.10–3.13 only) and potential rate-limiting or API stability issues with DuckDuckGo and Jina Reader, which the fact sheet…
This package is deprecated; install web-forager instead. It provides web research skills for AI agents, including DuckDuckGo search, page fetching, news monitoring, fact-checking, and article auditing.
Computes phonetic keys of strings using Soundex, NYSIIS, Metaphone, and Double Metaphone algorithms for fuzzy matching and phonetic indexing.
Client SDK for the ScrapeGraphAI managed API, enabling web scraping, structured data extraction, search, crawling, and monitoring through a cloud service with built-in LLM and browser rendering.
Programmatic access to Google Flights search results through reverse-engineered API calls, with CLI and MCP server interfaces for flight searches, date range queries, and multi-city itineraries.
Parses and crawls sitemaps in multiple formats (XML, RSS, Atom, plain text, Google News/Image) and extracts URLs efficiently without loading entire trees into memory.
Parses OData v4 filter strings and transpiles them to Django QuerySets, SQLAlchemy queries, or raw SQL for database filtering.
However, verify compatibility with your specific ORM versions before relying on it in production, since the last release was 777 days ago and no updates are forthcoming.
A CLI tool that auto-detects and downloads content from URLs using a plugin-based architecture, supporting multiple extraction methods (HTML, PDFs, screenshots, metadata, media) in a single command.
Python client library for the Rockset API, enabling programmatic creation, management, and querying of Rockset resources.
Scrape tweets, profiles, followers, and user data from Twitter/X using the platform's internal GraphQL API, authenticated with browser cookies instead of an official API key.
Query news articles and events from Event Registry's API using filters like keywords, concepts, sentiment, location, date, and category.
DocArray provides a Python data structure for representing, transmitting, storing, and retrieving multimodal data, with built-in support for tensors from NumPy, PyTorch, TensorFlow, and JAX.
Seltz is a Python SDK that provides web search and AI-powered answer generation for building search-augmented applications, with support for streaming and async operations.
However, verify the license status before use in proprietary projects, and confirm the service's production readiness and pricing model for your use case.
Python 3 client library for querying and retrieving real estate data from RETS (Real Estate Transaction Standard) Version 1.7.2 servers, including property listings, metadata, and associated media.
However, maintenance is dormant—last release was 668 days ago—so do not expect bug fixes, security patches, or support for new RETS features.
Jina is a framework for building and deploying AI services that communicate via gRPC, HTTP, and WebSockets, with built-in support for scaling, containerization, and cloud deployment.
Not recommended if you prefer minimal dependencies or need cutting-edge feature velocity.
Python bindings to libpostal, a C library for parsing and normalizing street addresses worldwide, supporting address expansion and component extraction.