Todas las etiquetas
Etiqueta del catálogo

#content-extraction

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

13 registros
Con la etiqueta «content-extraction»Ordenado por índice de salud
npm
89Excelenteíndice de salud
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7332↓ 492.7K/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
87Excelenteíndice de salud
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3201↓ 3522/mes22 jul 2026
Licencia propia22 jul 2026 · métricas 2.10.0
npm
86Excelenteíndice de salud
kepano/defuddle
Get the main content of any page as Markdown.
TypeScript★ 9191↓ 2.2M/mes28 ago 2026
MIT28 ago 2026 · métricas 2.10.0
Packagist
80Excelenteíndice de salud
j0k3r/php-readability
A fork of https://bitbucket.org/fivefilters/php-readability
PHP★ 184↓ 31.1K/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
Go · PyPI
80Excelenteíndice de salud
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 4220 jul 2026
MIT20 jul 2026 · métricas 2.10.0
71Buenoíndice de salud
extractus/article-extractor
To extract article from given URL
TypeScript · HTML★ 191028 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm
69Buenoíndice de salud
Thinkscape/agent-smart-fetch
Smarter, anti-bot resistant way to fetch stuff off the Internet
TypeScript★ 60↓ 5484/mes1 sept 2026
MIT1 sept 2026 · métricas 2.10.0
Go
65Buenoíndice de salud
dotcommander/defuddle
Go library and CLI for extracting web page content — articles, metadata, and clean text from any URL
Go★ 217 jul 2026
MIT17 jul 2026 · métricas 2.10.0
PyPI
56Moderadoíndice de salud
dondai1234/master-fetch
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
Python · HTML★ 33↓ 4672/mes18 jul 2026
MIT18 jul 2026 · métricas 2.10.0
npm
54Moderadoíndice de salud
valyuAI/valyu-js
The Official Valyu JavaScript SDK
TypeScript · JavaScript★ 2↓ 15.3K/mes17 jul 2026
Sin licencia17 jul 2026 · métricas 2.10.0
npm
53Moderadoíndice de salud
febbyRG/pdf-decomposer
A TypeScript Node.js library to parse all PDF page content (text, images, annotations, etc.) into JSON format.
TypeScript★ 2↓ 2074/mes31 jul 2026
Licencia propia31 jul 2026 · métricas 2.10.0
npm
51Moderadoíndice de salud
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/mes22 jul 2026
Sin licencia22 jul 2026 · métricas 2.10.0
PyPI · npm
39Débilíndice de salud
dondai44423/master-fetch
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
Python · HTML★ 430 jul 2026
MIT30 jul 2026 · métricas 2.10.0