Todas las etiquetas
Etiqueta del catálogo

#data-cleaning

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

6 registros
Con la etiqueta «data-cleaning»Ordenado por índice de salud
PyPI · npm
98Excepcionalíndice de salud
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
pyjanitor-devs/pyjanitor
Clean APIs for data cleaning. Python implementation of R package Janitor
Python★ 149913 ago 2026
MIT13 ago 2026 · métricas 2.10.0
PyPI
90Excelenteíndice de salud
unionai-oss/pandera
A light-weight, flexible, and expressive statistical data testing library
Python★ 444328 ago 2026
MIT28 ago 2026 · métricas 2.10.0
npm · PyPI
88Excelenteíndice de salud
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1316 sept 2026
MIT6 sept 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2527/mes16 ago 2026
Apache-2.016 ago 2026 · métricas 2.10.0
PyPI
50Moderadoíndice de salud
Kenzy-Zero/kenze
Big-file data prep that never runs out of memory - an interactive shell and CLI, powered by DuckDB.
Python★ 95 ago 2026
MIT5 ago 2026 · métricas 2.10.0