All tags
Catalogue tag

#html-parsing

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

6 records
Tagged “html-parsing”Ranked by health index
PyPI
94Exceptionalhealth index
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Python★ 72.6KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
npm
94Exceptionalhealth index
inikulin/parse5
HTML parsing/serialization toolset for Node.js. WHATWG HTML Living Standard (aka HTML5)-compliant.
TypeScript★ 3,918↓ 802M/moAug 4, 2026
MITAug 4, 2026 · metrics 2.10.0
Go
89Excellenthealth index
PuerkitoBio/goquery
A little like that j-thing, only in Go.
Go · Roff★ 15KAug 4, 2026
BSD-3-ClauseAug 4, 2026 · metrics 2.10.0
PyPI
89Excellenthealth index
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Python★ 1,067↓ 111.2K/moAug 22, 2026
MITAug 22, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
adbar/htmldate
Fast and robust date extraction from web pages, with Python or on the command-line
Python★ 155↓ 15M/moAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
miso-belica/jusText
Heuristic based boilerplate removal tool
Python★ 824↓ 12.3M/moAug 27, 2026
BSD-2-ClauseAug 27, 2026 · metrics 2.10.0