全部标签
目录标签

#data-quality

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

18 条记录
标签为“data-quality”按健康指数排序
PyPI · npm · Go +1
98卓越健康指数
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7,233↓ 841.1K/月2026年8月28日
Apache-2.02026年8月28日 · 指标 2.10.0
PyPI
98卓越健康指数
fivetran/great_expectations
Always know what to expect from your data.
Python★ 11.7K2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
PyPI · Maven · npm
98卓越健康指数
open-metadata/OpenMetadata
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
TypeScript · Java · Python★ 14.9K↓ 473.2K/月2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI · npm
98卓越健康指数
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K2026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
Go · Maven · npm
96卓越健康指数
treeverse/lakeFS
lakeFS - Data version control for your data lake | Git for data
Go★ 5,4852026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
PyPI · npm
90优秀健康指数
evidentlyai/evidently
Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Jupyter Notebook · Python★ 7,8002026年8月12日
Apache-2.02026年8月12日 · 指标 2.10.0
npm · PyPI
88优秀健康指数
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1312026年9月6日
MIT2026年9月6日 · 指标 2.10.0
PyPI
78良好健康指数
Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Python★ 13.7K↓ 18K/月2026年8月12日
MIT2026年8月12日 · 指标 2.10.0
PyPI
71良好健康指数
seadonggyun4/truthound
"Sniffs out bad data"
Python★ 18↓ 4,082/月2026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
Go
67良好健康指数
realdatadriven/etlx
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
Go★ 532026年9月5日
MIT2026年9月5日 · 指标 2.10.0
PyPI
60中等健康指数
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/月2026年7月19日
自定义许可证2026年7月19日 · 指标 2.10.0
PyPI
60中等健康指数
adidas/lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
Python★ 294↓ 3,910/月2026年8月18日
Apache-2.02026年8月18日 · 指标 2.10.0
Maven · PyPI
59中等健康指数
sparkutils/quality
A Quality Spark DQ and transformation Library
Scala★ 52026年9月5日
Apache-2.02026年9月5日 · 指标 2.10.0
PyPI
54中等健康指数
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2,527/月2026年8月16日
Apache-2.02026年8月16日 · 指标 2.10.0
Go · npm
54中等健康指数
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 12026年7月25日
Apache-2.02026年7月25日 · 指标 2.10.0
Hex
53中等健康指数
nshkrdotcom/json_remedy
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Elixir★ 33↓ 2,708/月2026年7月17日
MIT2026年7月17日 · 指标 2.10.0
Maven · PyPI
34存在风险健康指数
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 2,8282026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0
PyPI
19危急健康指数
datafold/data-diff
Compare tables within or across databases
Python★ 2,9882026年8月13日
MIT2026年8月13日 · 指标 2.10.0