全部标签
目录标签

#data-cleaning

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

6 条记录
标签为“data-cleaning”按健康指数排序
PyPI · npm
98卓越健康指数
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI
93卓越健康指数
pyjanitor-devs/pyjanitor
Clean APIs for data cleaning. Python implementation of R package Janitor
Python★ 1,4992026年8月13日
MIT2026年8月13日 · 指标 2.10.0
PyPI
90优秀健康指数
unionai-oss/pandera
A light-weight, flexible, and expressive statistical data testing library
Python★ 4,4432026年8月28日
MIT2026年8月28日 · 指标 2.10.0
npm · PyPI
88优秀健康指数
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1312026年9月6日
MIT2026年9月6日 · 指标 2.10.0
PyPI
54中等健康指数
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2,527/月2026年8月16日
Apache-2.02026年8月16日 · 指标 2.10.0
PyPI
50中等健康指数
Kenzy-Zero/kenze
Big-file data prep that never runs out of memory - an interactive shell and CLI, powered by DuckDB.
Python★ 92026年8月5日
MIT2026年8月5日 · 指标 2.10.0