All tags
Catalogue tag

#content-extraction

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

7 records
Tagged “content-extraction”Ranked by health index
npm
73Goodhealth index
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 1.13.0
Packagist
70Goodhealth index
j0k3r/php-readability
A fork of https://bitbucket.org/fivefilters/php-readability
PHP★ 184↓ 31.1K/moJul 20, 2026
Apache-2.0Jul 20, 2026 · metrics 1.13.0
Go · PyPI
68Moderatehealth index
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 42Jul 20, 2026
MITJul 20, 2026 · metrics 1.13.0
Go
59Moderatehealth index
dotcommander/defuddle
Go library and CLI for extracting web page content — articles, metadata, and clean text from any URL
Go★ 2Jul 17, 2026
MITJul 17, 2026 · metrics 1.13.0
PyPI
55Moderatehealth index
dondai1234/master-fetch
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
Python · HTML★ 33↓ 4,672/moJul 18, 2026
MITJul 18, 2026 · metrics 1.13.0
npm
52Moderatehealth index
valyuAI/valyu-js
The Official Valyu JavaScript SDK
TypeScript · JavaScript★ 2↓ 15.3K/moJul 17, 2026
No licenseJul 17, 2026 · metrics 1.13.0
npm
51Moderatehealth index
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/moJul 22, 2026
No licenseJul 22, 2026 · metrics 1.13.0