npm59Moderatehealth index

brandonkramer/pi-scraperPi extension for fast page scraping, recursive crawling, URL/site mapping, brand extraction, content diffing, PDF text extraction, and deterministic vertical extraction.
TypeScript · HTML★ 7↓ 607/moAug 28, 2026
Packagist57Moderatehealth index
PHP★ 324↓ 61.2K/moJul 28, 2026
PyPI57Moderatehealth index

codelucas/newspapernewspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
Python★ 15.1K↓ 788.5K/moAug 12, 2026
Go★ 70Jul 17, 2026
npm54Moderatehealth index
codepurse/SEOCOREEnterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health audits.
TypeScript★ 109Jul 17, 2026
PyPI54Moderatehealth index
tobybgy-lsd/web-agent-runtime-benchLocal-first failure diagnosis, auto collection, repair planning, AI handoff, and verification for Playwright, crawler, RPA, and agent workflows.
Python · JavaScript★ 1↓ 3,325/moJul 26, 2026
npm51Moderatehealth index
JavaScript★ 44↓ 560/moJul 22, 2026
Packagist · npm50Moderatehealth index
duzun/hQuery.phpAn extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/moJul 29, 2026
mjc-gh/virgoConvert a webpage into plaintext or markdown using the Chrome DevTool Protocol
Go · JavaScript★ 1Jul 17, 2026

quantmind-br/repodocsGo CLI for extracting websites, repositories, sitemaps, and package docs into structured Markdown
Go★ 0Aug 25, 2026
crates.io50Moderatehealth index
Rust★ 1↓ 2,668/moJul 16, 2026
npm · Go · PyPI48Weakhealth index
AlphaTechini/doc-fetchDynamic documentation fetching CLI that converts entire documentation sites to single markdown files for AI/LLM consumption
Svelte · Go★ 1↓ 2,547/moJul 27, 2026
Packagist42Weakhealth index
PHP★ 16↓ 2,779/moJul 15, 2026
Packagist41Weakhealth index
zorlan/skycaiji蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
PHP★ 2,078Jul 27, 2026
crates.io39Weakhealth index
Rust★ 4Jul 22, 2026
JavaScript★ 19↓ 4,152/moSep 2, 2026
C#★ 62Jul 31, 2026

scrapy/scrapyScrapy, a fast high-level web crawling & scraping framework for Python.
Python★ 63.8KAug 12, 2026
TypeScript★ 1↓ 2,234/moJul 25, 2026
Java · HTML★ 11.7KAug 27, 2026
JavaScript★ 9Jul 19, 2026
Synoppy/synoppy-goOfficial Go SDK for Synoppy — the web-data layer for AI agents. Read, crawl, map, extract, classify & enrich any website on one key. Standard library only.
Go★ 1Jul 28, 2026
Java★ 11.4KAug 27, 2026
PyPI22At Riskhealth index
Python · HTML★ 1↓ 4,246/moAug 1, 2026
npm · PyPI20At Riskhealth index
Python · JavaScript★ 9Jul 27, 2026
karthikuj/sasoriSasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
JavaScript★ 145↓ 78/moAug 4, 2026