
apify/crawleeCrawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/月2026年8月5日

apify/crawlee-pythonCrawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,4262026年8月12日

apify/apify-client-pythonApify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/月2026年8月27日

D4Vinci/Scrapling🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
Python★ 29K2026年8月5日
apify/apify-sdk-pythonApify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Packagist · Hex · crates.io +294卓越健康指数
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日

plabayo/ramamodular service framework to move and transform network packets
Rust★ 1,190↓ 730.4K/月2026年9月5日
HTML★ 623↓ 25K/月2026年8月18日

adbar/trafilaturaPython & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/月2026年8月28日
TypeScript · JavaScript★ 2,566↓ 2.2M/月2026年8月13日
Python · HTML★ 16.5K2026年8月5日

soxoj/maigret🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Python · HTML★ 36.1K↓ 99.4K/月2026年8月5日

soxoj/socid-extractor⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,067↓ 111.2K/月2026年8月22日

feder-cr/invisible_playwrightFree antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/月2026年9月4日
KnockOutEZ/wigoloThe go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
C++ · Python · JavaScript★ 11.7K↓ 851.4K/月2026年9月6日

jo-inc/camofox-browserStealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8,924↓ 243.1K/月2026年8月28日
Rust★ 2,669↓ 42K/月2026年8月22日

getmaxun/maxun🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日

ScrapeGraphAI/scrapegraph-pyOfficial Python SDK for the ScrapeGraph AI API. Smart scraping, search, crawling, markdownify, agentic browser automation, scheduled jobs, and structured data extraction
Jupyter Notebook★ 86↓ 212.9K/月2026年9月7日
achiya-automation/safari-mcpNative Safari browser automation for AI agents. 80 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chrome DevTools MCP with 40-60% less CPU/heat on Apple Silicon.
JavaScript★ 151↓ 6,053/月2026年7月17日
blisspixel/primrTurn any company URL into a strategic intelligence brief. Adaptive scraping + AI-powered research and synthesis.
Python★ 3↓ 7,380/月2026年7月19日
JavaScript★ 357↓ 20.3K/月2026年8月3日
ArchiveBox/abx-plugins🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/月2026年7月27日
PyPI · crates.io · npm81优秀健康指数

bug-ops/scrape-rs🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/月2026年9月5日
ArchiveBox/abx-dl⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/月2026年7月19日

CloakHQ/CloakBrowserStealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/月2026年8月5日

brightdata/brightdata-mcpA powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/月2026年8月28日
Python★ 2,381↓ 1.9M/月2026年8月3日