All tags
Catalogue tag

#web-crawling

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

9 records
Tagged “web-crawling”Ranked by health index
npm
98Exceptionalhealth index
apify/crawlee
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
Apache-2.0Aug 5, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
97Exceptionalhealth index
apify/apify-client-python
Apify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 2.10.0
PyPI
91Excellenthealth index
seleniumbase/SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
Python★ 13K↓ 3M/moAug 27, 2026
MITAug 27, 2026 · metrics 2.10.0
npm
80Excellenthealth index
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
npm
67Goodhealth index
superagents-lab/search1api-mcp
Official Search1API MCP server for web search, news, crawling, sitemaps, and trends—hosted with OAuth 2.1 or local via npm.
TypeScript · JavaScript★ 173↓ 2,722/moAug 24, 2026
MITAug 24, 2026 · metrics 2.10.0
npm · crates.io · Go +1
63Moderatehealth index
spider-rs/spider-clients
Python, Javascript, and Rust libraries for the Spider Cloud API.
Rust · Python · Go★ 26↓ 16.3K/moJul 16, 2026
MITJul 16, 2026 · metrics 2.10.0
npm
47Weakhealth index
NovadaLabs/novada-mcp
One MCP server for all web data — search, scrape, crawl, proxy, and AI research in a single npx install. Works with Claude, Cursor, and any MCP client.
TypeScript · HTML★ 2↓ 5,804/moJul 16, 2026
No licenseJul 16, 2026 · metrics 2.10.0