npm98Exceptionalhealth index

apify/crawleeCrawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
TypeScript · MDX★ 25.2K↓ 6M/moAug 5, 2026
PyPI · npm98Exceptionalhealth index

apify/crawlee-pythonCrawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python · MDX★ 9,426Aug 12, 2026
PyPI · npm97Exceptionalhealth index

apify/apify-client-pythonApify API client for Python—Programmatically run Actors, manage and stream data from storages (datasets, key-value stores, request queues), schedule and monitor runs, and access the full Apify platform API. Sync and async interfaces with automatic retries and pagination.
Python★ 97↓ 2.8M/moAug 27, 2026
PyPI94Exceptionalhealth index

D4Vinci/Scrapling🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
PyPI94Exceptionalhealth index
Python★ 29KAug 5, 2026
PyPI · npm94Exceptionalhealth index
apify/apify-sdk-pythonApify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 173Jul 20, 2026
Packagist · Hex · crates.io +294Exceptionalhealth index
TypeScript · Python★ 161K↓ 3,278/moAug 4, 2026
crates.io94Exceptionalhealth index

plabayo/ramamodular service framework to move and transform network packets
Rust★ 1,190↓ 730.4K/moSep 5, 2026
PyPI93Exceptionalhealth index
HTML★ 623↓ 25K/moAug 18, 2026
PyPI91Excellenthealth index

adbar/trafilaturaPython & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Python★ 6,718↓ 14M/moAug 28, 2026
PyPI · npm90Excellenthealth index
TypeScript · JavaScript★ 2,566↓ 2.2M/moAug 13, 2026
PyPI90Excellenthealth index
Python · HTML★ 16.5KAug 5, 2026
PyPI89Excellenthealth index

soxoj/maigret🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Python · HTML★ 36.1K↓ 99.4K/moAug 5, 2026
PyPI89Excellenthealth index

soxoj/socid-extractor⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,067↓ 111.2K/moAug 22, 2026
PyPI88Excellenthealth index

feder-cr/invisible_playwrightFree antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
Python★ 1,974↓ 15.7K/moSep 4, 2026
npm87Excellenthealth index
KnockOutEZ/wigoloThe go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
PyPI · npm87Excellenthealth index
C++ · Python · JavaScript★ 11.7K↓ 851.4K/moSep 6, 2026
npm87Excellenthealth index

jo-inc/camofox-browserStealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
JavaScript★ 8,924↓ 243.1K/moAug 28, 2026
crates.io87Excellenthealth index
Rust★ 2,669↓ 42K/moAug 22, 2026
npm86Excellenthealth index

getmaxun/maxun🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/moAug 5, 2026
npm83Excellenthealth index
achiya-automation/safari-mcpNative Safari browser automation for AI agents. 80 tools via AppleScript — zero overhead, keeps logins, runs silently in background. Drop-in alternative to Chrome DevTools MCP with 40-60% less CPU/heat on Apple Silicon.
JavaScript★ 151↓ 6,053/moJul 17, 2026
PyPI83Excellenthealth index
blisspixel/primrTurn any company URL into a strategic intelligence brief. Adaptive scraping + AI-powered research and synthesis.
Python★ 3↓ 7,380/moJul 19, 2026
npm83Excellenthealth index
JavaScript★ 357↓ 20.3K/moAug 3, 2026
PyPI81Excellenthealth index
ArchiveBox/abx-plugins🧩 Plugins and extractors that ArchiveBox + abx-dl use: chrome, ytdlp, wget, singlefile, readability, forum-dl, gallery-dl, papers-dl, and more...
Python · JavaScript★ 8↓ 41.2K/moJul 27, 2026
PyPI · crates.io · npm81Excellenthealth index

bug-ops/scrape-rs🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/moSep 5, 2026
PyPI · npm80Excellenthealth index
ArchiveBox/abx-dl⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
Python★ 130↓ 9,160/moJul 19, 2026
PyPI · npm80Excellenthealth index

CloakHQ/CloakBrowserStealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python · C# · TypeScript★ 29.6K↓ 1M/moAug 5, 2026
npm80Excellenthealth index

brightdata/brightdata-mcpA powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/moAug 28, 2026
Python★ 2,381↓ 1.9M/moAug 3, 2026
Packagist · npm78Goodhealth index
playwright-php/playwrightPlaywright PHP library for browser automation: navigation, E2E tests, assertions, screenshots, and so much more!
PHP★ 204↓ 12.7K/moJul 18, 2026