全部标签
目录标签

#html-parser

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

15 条记录
标签为“html-parser”按健康指数排序
npm
94卓越健康指数
inikulin/parse5
HTML parsing/serialization toolset for Node.js. WHATWG HTML Living Standard (aka HTML5)-compliant.
TypeScript★ 3,918↓ 802M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
npm
91优秀健康指数
fb55/htmlparser2
The fast & forgiving HTML and XML parser
TypeScript★ 4,785↓ 346M/月2026年8月4日
MIT2026年8月4日 · 指标 2.10.0
PyPI · crates.io · npm
81优秀健康指数
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
Hex
81优秀健康指数
philss/floki
Floki is a simple HTML parser that enables search for nodes using CSS selectors.
Elixir★ 2,145↓ 423.3K/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
Packagist
71良好健康指数
voku/simple_html_dom
📜 Modern Simple HTML DOM Parser for PHP
PHP★ 900↓ 226.4K/月2026年7月15日
MIT2026年7月15日 · 指标 2.10.0
crates.io · npm · PyPI
69良好健康指数
nabbisen/mdka-rs
A HTML to Markdown (MD) converter balances conversion quality with runtime efficiency.
Rust · JavaScript · Python★ 54↓ 5,080/月2026年8月3日
Apache-2.02026年8月3日 · 指标 2.10.0
Go
65良好健康指数
dotcommander/defuddle
Go library and CLI for extracting web page content — articles, metadata, and clean text from any URL
Go★ 22026年7月17日
MIT2026年7月17日 · 指标 2.10.0
npm
63中等健康指数
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/月2026年9月5日
MIT2026年9月5日 · 指标 2.10.0
NuGet
62中等健康指数
zzzprojects/html-agility-pack
Html Agility Pack (HAP) is a free and open-source HTML parser written in C# to read/write DOM and supports plain XPATH or XSLT. It is a .NET code library that allows you to parse "out of the web" HTML files.
C# · HTML★ 2,8462026年8月22日
MIT2026年8月22日 · 指标 2.10.0
Hex
60中等健康指数
rusterlium/html5ever_elixir
NIF wrapper of html5ever using Rustler
HTML · Rust · Elixir★ 87↓ 7,019/月2026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
PyPI
54中等健康指数
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/月2026年8月27日
BSD-2-Clause2026年8月27日 · 指标 2.10.0
Packagist · npm
50中等健康指数
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/月2026年7月29日
MIT2026年7月29日 · 指标 2.10.0
NuGet
35薄弱健康指数
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 622026年7月31日
自定义许可证2026年7月31日 · 指标 2.10.0
Go
11危急健康指数
andreychh/tgen
Turns the Telegram Bot API HTML documentation into ready-to-use API bindings.
Python · Go★ 22026年7月20日
MIT2026年7月20日 · 指标 2.10.0
Go
7危急健康指数
wmentor/html
HTML data fetcher
Go★ 12026年9月3日
MIT2026年9月3日 · 指标 2.10.0