All tags
Catalogue tag

#data-quality

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

18 records
Tagged “data-quality”Ranked by health index
PyPI · npm · Go +1
98Exceptionalhealth index
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7,233↓ 841.1K/moAug 28, 2026
Apache-2.0Aug 28, 2026 · metrics 2.10.0
PyPI
98Exceptionalhealth index
fivetran/great_expectations
Always know what to expect from your data.
Python★ 11.7KAug 27, 2026
Apache-2.0Aug 27, 2026 · metrics 2.10.0
PyPI · Maven · npm
98Exceptionalhealth index
open-metadata/OpenMetadata
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
TypeScript · Java · Python★ 14.9K↓ 473.2K/moAug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
98Exceptionalhealth index
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11KAug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
Go · Maven · npm
96Exceptionalhealth index
treeverse/lakeFS
lakeFS - Data version control for your data lake | Git for data
Go★ 5,485Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
PyPI · npm
90Excellenthealth index
evidentlyai/evidently
Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Jupyter Notebook · Python★ 7,800Aug 12, 2026
Apache-2.0Aug 12, 2026 · metrics 2.10.0
npm · PyPI
88Excellenthealth index
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 131Sep 6, 2026
MITSep 6, 2026 · metrics 2.10.0
PyPI
78Goodhealth index
Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Python★ 13.7K↓ 18K/moAug 12, 2026
MITAug 12, 2026 · metrics 2.10.0
PyPI
71Goodhealth index
seadonggyun4/truthound
"Sniffs out bad data"
Python★ 18↓ 4,082/moSep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
Go
67Goodhealth index
realdatadriven/etlx
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
Go★ 53Sep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/moJul 19, 2026
Custom licenseJul 19, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
adidas/lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
Python★ 294↓ 3,910/moAug 18, 2026
Apache-2.0Aug 18, 2026 · metrics 2.10.0
Maven · PyPI
59Moderatehealth index
sparkutils/quality
A Quality Spark DQ and transformation Library
Scala★ 5Sep 5, 2026
Apache-2.0Sep 5, 2026 · metrics 2.10.0
PyPI
54Moderatehealth index
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2,527/moAug 16, 2026
Apache-2.0Aug 16, 2026 · metrics 2.10.0
Go · npm
54Moderatehealth index
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 1Jul 25, 2026
Apache-2.0Jul 25, 2026 · metrics 2.10.0
Hex
53Moderatehealth index
nshkrdotcom/json_remedy
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Elixir★ 33↓ 2,708/moJul 17, 2026
MITJul 17, 2026 · metrics 2.10.0
Maven · PyPI
34At Riskhealth index
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 2,828Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
19Criticalhealth index
datafold/data-diff
Compare tables within or across databases
Python★ 2,988Aug 13, 2026
MITAug 13, 2026 · metrics 2.10.0