Alle Tags
Katalog-Tag

#data-quality

Alle Repositories im öffentlichen Register, die dieses Tag tragen — aus ihren GitHub-Topics oder den von ihren Paket-Registries veröffentlichten Schlagwörtern. Die Gesundheit wird nach derselben versionierten Methodik gemessen wie im übrigen Register.

18 Einträge
Getaggt als „data-quality“Geordnet nach Gesundheitsindex
PyPI · npm · Go +1
98AußergewöhnlichGesundheitsindex
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7.233↓ 841.1K/Monat28. Aug. 2026
Apache-2.028. Aug. 2026 · Metriken 2.10.0
PyPI
98AußergewöhnlichGesundheitsindex
fivetran/great_expectations
Always know what to expect from your data.
Python★ 11.7K27. Aug. 2026
Apache-2.027. Aug. 2026 · Metriken 2.10.0
PyPI · Maven · npm
98AußergewöhnlichGesundheitsindex
open-metadata/OpenMetadata
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
TypeScript · Java · Python★ 14.9K↓ 473.2K/Monat12. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI · npm
98AußergewöhnlichGesundheitsindex
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K12. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
Go · Maven · npm
96AußergewöhnlichGesundheitsindex
treeverse/lakeFS
lakeFS - Data version control for your data lake | Git for data
Go★ 5.48512. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
PyPI · npm
90ExzellentGesundheitsindex
evidentlyai/evidently
Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Jupyter Notebook · Python★ 7.80012. Aug. 2026
Apache-2.012. Aug. 2026 · Metriken 2.10.0
npm · PyPI
88ExzellentGesundheitsindex
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1316. Sept. 2026
MIT6. Sept. 2026 · Metriken 2.10.0
PyPI
78GutGesundheitsindex
Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Python★ 13.7K↓ 18K/Monat12. Aug. 2026
MIT12. Aug. 2026 · Metriken 2.10.0
PyPI
71GutGesundheitsindex
seadonggyun4/truthound
"Sniffs out bad data"
Python★ 18↓ 4.082/Monat5. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
Go
67GutGesundheitsindex
realdatadriven/etlx
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
Go★ 535. Sept. 2026
MIT5. Sept. 2026 · Metriken 2.10.0
PyPI
60MittelGesundheitsindex
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/Monat19. Juli 2026
Eigene Lizenz19. Juli 2026 · Metriken 2.10.0
PyPI
60MittelGesundheitsindex
adidas/lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
Python★ 294↓ 3.910/Monat18. Aug. 2026
Apache-2.018. Aug. 2026 · Metriken 2.10.0
Maven · PyPI
59MittelGesundheitsindex
sparkutils/quality
A Quality Spark DQ and transformation Library
Scala★ 55. Sept. 2026
Apache-2.05. Sept. 2026 · Metriken 2.10.0
PyPI
54MittelGesundheitsindex
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2.527/Monat16. Aug. 2026
Apache-2.016. Aug. 2026 · Metriken 2.10.0
Go · npm
54MittelGesundheitsindex
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 125. Juli 2026
Apache-2.025. Juli 2026 · Metriken 2.10.0
Hex
53MittelGesundheitsindex
nshkrdotcom/json_remedy
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Elixir★ 33↓ 2.708/Monat17. Juli 2026
MIT17. Juli 2026 · Metriken 2.10.0
Maven · PyPI
34GefährdetGesundheitsindex
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 2.82821. Juli 2026
Apache-2.021. Juli 2026 · Metriken 2.10.0
PyPI
19KritischGesundheitsindex
datafold/data-diff
Compare tables within or across databases
Python★ 2.98813. Aug. 2026
MIT13. Aug. 2026 · Metriken 2.10.0