全部标签
目录标签

#data-extraction

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

35 条记录
标签为“data-extraction”按健康指数排序
PyPI
94卓越健康指数
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6K2026年8月4日
BSD-3-Clause2026年8月4日 · 指标 2.10.0
PyPI
94卓越健康指数
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
Python★ 29K2026年8月5日
MIT2026年8月5日 · 指标 2.10.0
PyPI · npm
94卓越健康指数
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Python · MDX★ 1732026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
Packagist · Hex · crates.io +2
94卓越健康指数
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
TypeScript · Python★ 161K↓ 3,278/月2026年8月4日
AGPL-3.02026年8月4日 · 指标 2.10.0
crates.io · PyPI · npm +5
93卓越健康指数
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Rust★ 1,015↓ 391.7K/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
PyPI
89优秀健康指数
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,067↓ 111.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm · PyPI · crates.io
87优秀健康指数
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
Rust★ 518↓ 9,070/月2026年8月3日
AGPL-3.02026年8月3日 · 指标 2.10.0
crates.io · PyPI · npm +1
87优秀健康指数
yfedoseev/office_oxide
The fastest Office document library for Python, Rust, Go, JS/TS, C# and WASM. DOCX, XLSX, PPTX, DOC, XLS, PPT. Up to 100× faster than python-docx/openpyxl/python-pptx. 100% pass rate on valid Office files. MIT/Apache-2.0.
Rust★ 108↓ 161.1K/月2026年8月22日
Apache-2.02026年8月22日 · 指标 2.10.0
npm
86优秀健康指数
getmaxun/maxun
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
TypeScript★ 17K↓ 952/月2026年8月5日
AGPL-3.02026年8月5日 · 指标 2.10.0
crates.io
84优秀健康指数
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
Rust★ 185↓ 15.2K/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
Packagist · PyPI
84优秀健康指数
cognesy/instructor-php
Unified LLM API, structured data outputs with LLMs, and agent SDK - in PHP
PHP★ 325↓ 5,250/月2026年7月20日
MIT2026年7月20日 · 指标 2.10.0
PyPI
83优秀健康指数
thinh-vu/vnstock
A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone
Python★ 1,3582026年7月28日
自定义许可证2026年7月28日 · 指标 2.10.0
npm
81优秀健康指数
Xquik-dev/tweetclaw
OpenClaw plugin to search tweets, search replies, post tweets, export followers, manage media, monitor X/Twitter, and run giveaway draws via Xquik. Not affiliated with X Corp.
TypeScript · JavaScript★ 89↓ 1,243/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
PyPI
81优秀健康指数
Xquik-dev/x-twitter-scraper-python
Python REST SDK for Xquik tweet search, profiles, followers, media, publishing, and webhooks. Not affiliated with X Corp.
Python★ 3↓ 710/月2026年7月26日
Apache-2.02026年7月26日 · 指标 2.10.0
npm
80优秀健康指数
brightdata/brightdata-mcp
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
JavaScript★ 2,614↓ 28.7K/月2026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm · crates.io
75良好健康指数
0xMassi/webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Rust★ 2,066↓ 424/月2026年7月26日
AGPL-3.02026年7月26日 · 指标 2.10.0
npm
73良好健康指数
firecrawl/cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
TypeScript · JavaScript★ 539↓ 78.1K/月2026年7月26日
无许可证2026年7月26日 · 指标 2.10.0
Go
73良好健康指数
xquik-dev/terraform-provider-x-twitter-scraper
Terraform provider for Xquik Twitter automation: manage monitors, webhooks, API keys, tweet actions, follower exports, draws and extraction jobs as infrastructure. Not affiliated with X Corp.
Go★ 02026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
71良好健康指数
extractus/article-extractor
To extract article from given URL
TypeScript · HTML★ 1,9102026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm
71良好健康指数
scrape-badger/scrapebadger-node
Official Node.js SDK for ScrapeBadger - Async web scraping APIs for Twitter and more
TypeScript★ 2↓ 3,672/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
Go
69良好健康指数
Xquik-dev/x-twitter-scraper-go
Go REST SDK for Xquik: search posts, read profiles, export followers, download media, monitor accounts, publish, reply, and use webhooks. Not affiliated with X Corp.
Go★ 02026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Packagist · Go · crates.io +2
67良好健康指数
kreuzberg-dev/kreuzberg-lts
Kreuzberg v4 LTS — long-term support for the v4 line (legacy; superseded by xberg for v5+). MIT-licensed.
Rust · HTML★ 10↓ 1/月2026年8月22日
MIT2026年8月22日 · 指标 2.10.0
npm
67良好健康指数
xquik-dev/x-twitter-scraper
Xquik agent Skill for REST, OAuth-first MCP, webhooks, exports, tweet search, follower data, monitoring, and confirmation-gated publishing. Not affiliated with X Corp.
JavaScript★ 155↓ 1,004/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
Go
65良好健康指数
xquik-dev/x-twitter-scraper-cli
Command-line REST client for Xquik: search posts, read profiles, export followers, download media, publish, reply, and manage webhooks. Not affiliated with X Corp.
Go★ 02026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
npm · RubyGems
63中等健康指数
LeonTing1010/tap
Capture a logged-in browser task once — replay it forever at zero LLM tokens. Local-first browser-automation MCP for Claude Code, Cursor & any MCP host; credentials never leave your machine.
JavaScript · TypeScript★ 122026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm · PyPI
63中等健康指数
ShZhao27208/Aut_Sci_Write
Academic research skills suite for AI Agent — literature search/download (WoS+Elsevier+Springer), PDF extraction, figure cropping, review writing, Zotero sync, and PPT/Html generation.
Python★ 165↓ 1,096/月2026年7月20日
自定义许可证2026年7月20日 · 指标 2.10.0
npm
63中等健康指数
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
PyPI
62中等健康指数
SukramJ/openccu-data
Data-extraction pipeline and source-of-truth distribution for Homematic CCU configuration metadata — parses OCCU/RaspberryMatic TCL & JS into compact, typed JSON artifacts (easymodes, translations, profiles).
Python★ 0↓ 374.8K/月2026年7月19日
MIT2026年7月19日 · 指标 2.10.0
PyPI
56中等健康指数
ivan-loh/messy-xlsx
Python library for parsing Excel files with structure detection and normalization
Python★ 22026年8月30日
MIT2026年8月30日 · 指标 2.10.0
PyPI · npm
48薄弱健康指数
0m3rF/ELM-tool
Extract-Mask-Load tool between databases.
Python★ 0↓ 39/月2026年8月15日
GPL-3.02026年8月15日 · 指标 2.10.0