全部标签
目录标签

#content-extraction

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

13 条记录
标签为“content-extraction”按健康指数排序
npm
89优秀健康指数
firecrawl/firecrawl-mcp-server
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
TypeScript · JavaScript★ 7,332↓ 492.7K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm
87优秀健康指数
KnockOutEZ/wigolo
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
TypeScript★ 3,201↓ 3,522/月2026年7月22日
自定义许可证2026年7月22日 · 指标 2.10.0
npm
86优秀健康指数
kepano/defuddle
Get the main content of any page as Markdown.
TypeScript★ 9,191↓ 2.2M/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
Packagist
80优秀健康指数
j0k3r/php-readability
A fork of https://bitbucket.org/fivefilters/php-readability
PHP★ 184↓ 31.1K/月2026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
Go · PyPI
80优秀健康指数
zoharbabin/web-researcher-mcp
The AI research assistant that cites real sources honestly — and searches the web. Your AI research assistant that cites real sources and stays honest. Works with Claude, Cursor, any MCP client.
Go★ 422026年7月20日
MIT2026年7月20日 · 指标 2.10.0
71良好健康指数
extractus/article-extractor
To extract article from given URL
TypeScript · HTML★ 1,9102026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm
69良好健康指数
Thinkscape/agent-smart-fetch
Smarter, anti-bot resistant way to fetch stuff off the Internet
TypeScript★ 60↓ 5,484/月2026年9月1日
MIT2026年9月1日 · 指标 2.10.0
Go
65良好健康指数
dotcommander/defuddle
Go library and CLI for extracting web page content — articles, metadata, and clean text from any URL
Go★ 22026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
56中等健康指数
dondai1234/master-fetch
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
Python · HTML★ 33↓ 4,672/月2026年7月18日
MIT2026年7月18日 · 指标 2.10.0
npm
54中等健康指数
valyuAI/valyu-js
The Official Valyu JavaScript SDK
TypeScript · JavaScript★ 2↓ 15.3K/月2026年7月17日
无许可证2026年7月17日 · 指标 2.10.0
npm
53中等健康指数
febbyRG/pdf-decomposer
A TypeScript Node.js library to parse all PDF page content (text, images, annotations, etc.) into JSON format.
TypeScript★ 2↓ 2,074/月2026年7月31日
自定义许可证2026年7月31日 · 指标 2.10.0
npm
51中等健康指数
webscraping-ai/webscraping-ai-mcp-server
A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.
JavaScript★ 44↓ 560/月2026年7月22日
无许可证2026年7月22日 · 指标 2.10.0
PyPI · npm
39薄弱健康指数
dondai44423/master-fetch
MCP server for web fetching with Cloudflare bypass, Trafilatura extraction, and smart routing. Free, self-hosted, no API keys.
Python · HTML★ 42026年7月30日
MIT2026年7月30日 · 指标 2.10.0