Todas las etiquetas
Etiqueta del catálogo

#data-quality

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

18 registros
Con la etiqueta «data-quality»Ordenado por índice de salud
PyPI · npm · Go +1
98Excepcionalíndice de salud
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7233↓ 841.1K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
98Excepcionalíndice de salud
fivetran/great_expectations
Always know what to expect from your data.
Python★ 11.7K27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
PyPI · Maven · npm
98Excepcionalíndice de salud
open-metadata/OpenMetadata
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
TypeScript · Java · Python★ 14.9K↓ 473.2K/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
Go · Maven · npm
96Excepcionalíndice de salud
treeverse/lakeFS
lakeFS - Data version control for your data lake | Git for data
Go★ 548512 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · npm
90Excelenteíndice de salud
evidentlyai/evidently
Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Jupyter Notebook · Python★ 780012 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
npm · PyPI
88Excelenteíndice de salud
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1316 sept 2026
MIT6 sept 2026 · métricas 2.10.0
PyPI
78Buenoíndice de salud
Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Python★ 13.7K↓ 18K/mes12 ago 2026
MIT12 ago 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
seadonggyun4/truthound
"Sniffs out bad data"
Python★ 18↓ 4082/mes5 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
Go
67Buenoíndice de salud
realdatadriven/etlx
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
Go★ 535 sept 2026
MIT5 sept 2026 · métricas 2.10.0
PyPI
60Moderadoíndice de salud
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/mes19 jul 2026
Licencia propia19 jul 2026 · métricas 2.10.0
PyPI
60Moderadoíndice de salud
adidas/lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
Python★ 294↓ 3910/mes18 ago 2026
Apache-2.018 ago 2026 · métricas 2.10.0
Maven · PyPI
59Moderadoíndice de salud
sparkutils/quality
A Quality Spark DQ and transformation Library
Scala★ 55 sept 2026
Apache-2.05 sept 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2527/mes16 ago 2026
Apache-2.016 ago 2026 · métricas 2.10.0
Go · npm
54Moderadoíndice de salud
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 125 jul 2026
Apache-2.025 jul 2026 · métricas 2.10.0
Hex
53Moderadoíndice de salud
nshkrdotcom/json_remedy
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Elixir★ 33↓ 2708/mes17 jul 2026
MIT17 jul 2026 · métricas 2.10.0
Maven · PyPI
34En riesgoíndice de salud
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 282821 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
19Críticoíndice de salud
datafold/data-diff
Compare tables within or across databases
Python★ 298813 ago 2026
MIT13 ago 2026 · métricas 2.10.0