All tags
Catalogue tag

#html-parser

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

15 records
Tagged “html-parser”Ranked by health index
npm
94Exceptionalhealth index
inikulin/parse5
HTML parsing/serialization toolset for Node.js. WHATWG HTML Living Standard (aka HTML5)-compliant.
TypeScript★ 3,918↓ 802M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
npm
91Excellenthealth index
fb55/htmlparser2
The fast & forgiving HTML and XML parser
TypeScript★ 4,785↓ 346M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
PyPI · crates.io · npm
81Excellenthealth index
bug-ops/scrape-rs
🦀 High-performance HTML parsing library. Rust core with native bindings for Python, Node.js & WASM. SIMD-accelerated, memory-safe, consistent API everywhere.
Rust★ 10↓ 6,257/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
Hex
81Excellenthealth index
philss/floki
Floki is a simple HTML parser that enables search for nodes using CSS selectors.
Elixir★ 2,145↓ 423.3K/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
Packagist
71Goodhealth index
voku/simple_html_dom
📜 Modern Simple HTML DOM Parser for PHP
PHP★ 900↓ 226.4K/moJul 15, 2026
MITJul 15, 2026 · metrics 2.10.0
crates.io · npm · PyPI
69Goodhealth index
nabbisen/mdka-rs
A HTML to Markdown (MD) converter balances conversion quality with runtime efficiency.
Rust · JavaScript · Python★ 54↓ 5,080/moAug 3, 2026
Apache-2.0Aug 3, 2026 · metrics 2.10.0
Go
65Goodhealth index
dotcommander/defuddle
Go library and CLI for extracting web page content — articles, metadata, and clean text from any URL
Go★ 2Jul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
npm
63Moderatehealth index
mysleekdesigns/crawlforge-mcp
28 MCP tools that give Claude, Cursor and any MCP client the live web — scrape, crawl, search, real Google rank, change tracking, document parsing, plus an autonomous agent that researches from a plain-English prompt with no URLs. Clean Markdown and schema-validated JSON, not raw HTML. Local-Ollama extraction by default. MIT, 1,000 free credits.
JavaScript★ 2↓ 5,179/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
NuGet
62Moderatehealth index
zzzprojects/html-agility-pack
Html Agility Pack (HAP) is a free and open-source HTML parser written in C# to read/write DOM and supports plain XPATH or XSLT. It is a .NET code library that allows you to parse "out of the web" HTML files.
C# · HTML★ 2,846Aug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
Hex
60Moderatehealth index
rusterlium/html5ever_elixir
NIF wrapper of html5ever using Rustler
HTML · Rust · Elixir★ 87↓ 7,019/moJul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/moAug 27, 2026
BSD-2-ClauseAug 27, 2026 · metrics 2.10.0
Packagist · npm
50Moderatehealth index
duzun/hQuery.php
An extremely fast web scraper that parses megabytes of invalid HTML in a blink of an eye. PHP5.3+, no dependencies.
PHP★ 360↓ 12.6K/moJul 29, 2026
MITJul 29, 2026 · metrics 2.10.0
NuGet
35Weakhealth index
SoftCircuits/HtmlMonkey
Lightweight HTML/XML parser written in C#.
C#★ 62Jul 31, 2026
Custom licenseJul 31, 2026 · metrics 2.10.0
Go
11Criticalhealth index
andreychh/tgen
Turns the Telegram Bot API HTML documentation into ready-to-use API bindings.
Python · Go★ 2Jul 20, 2026
MITJul 20, 2026 · metrics 2.10.0
Go
7Criticalhealth index
wmentor/html
HTML data fetcher
Go★ 1Sep 3, 2026
MITSep 3, 2026 · metrics 2.10.0