Усі теги
Тег каталогу

#data-quality

Усі репозиторії публічного реєстру з цим тегом — із тем GitHub або ключових слів, які публікують їхні реєстри пакетів. Здоров'я вимірюється за тією ж версіонованою методологією, що й решта реєстру.

18 записів
З тегом «data-quality»Упорядковано за індексом здоров'я
PyPI · npm · Go +1
98Винятковийіндекс здоров'я
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7 233↓ 841.1K/міс28 серп. 2026 р.
Apache-2.028 серп. 2026 р. · метрики 2.10.0
PyPI
98Винятковийіндекс здоров'я
fivetran/great_expectations
Always know what to expect from your data.
Python★ 11.7K27 серп. 2026 р.
Apache-2.027 серп. 2026 р. · метрики 2.10.0
PyPI · Maven · npm
98Винятковийіндекс здоров'я
open-metadata/OpenMetadata
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
TypeScript · Java · Python★ 14.9K↓ 473.2K/міс12 серп. 2026 р.
Apache-2.012 серп. 2026 р. · метрики 2.10.0
PyPI · npm
98Винятковийіндекс здоров'я
voxel51/fiftyone
Refine high-quality datasets and visual AI models
TypeScript · Python★ 11K12 серп. 2026 р.
Apache-2.012 серп. 2026 р. · метрики 2.10.0
Go · Maven · npm
96Винятковийіндекс здоров'я
treeverse/lakeFS
lakeFS - Data version control for your data lake | Git for data
Go★ 5 48512 серп. 2026 р.
Apache-2.012 серп. 2026 р. · метрики 2.10.0
PyPI · npm
90Відміннийіндекс здоров'я
evidentlyai/evidently
Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Jupyter Notebook · Python★ 7 80012 серп. 2026 р.
Apache-2.012 серп. 2026 р. · метрики 2.10.0
npm · PyPI
88Відміннийіндекс здоров'я
benseverndev-oss/goldenmatch
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Python · TypeScript★ 1316 вер. 2026 р.
MIT6 вер. 2026 р. · метрики 2.10.0
PyPI
78Добрийіндекс здоров'я
Data-Centric-AI-Community/fg-data-profiling
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Python★ 13.7K↓ 18K/міс12 серп. 2026 р.
MIT12 серп. 2026 р. · метрики 2.10.0
PyPI
71Добрийіндекс здоров'я
seadonggyun4/truthound
"Sniffs out bad data"
Python★ 18↓ 4 082/міс5 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
Go
67Добрийіндекс здоров'я
realdatadriven/etlx
ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.
Go★ 535 вер. 2026 р.
MIT5 вер. 2026 р. · метрики 2.10.0
PyPI
60Помірнийіндекс здоров'я
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/міс19 лип. 2026 р.
Власна ліцензія19 лип. 2026 р. · метрики 2.10.0
PyPI
60Помірнийіндекс здоров'я
adidas/lakehouse-engine
The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.
Python★ 294↓ 3 910/міс18 серп. 2026 р.
Apache-2.018 серп. 2026 р. · метрики 2.10.0
Maven · PyPI
59Помірнийіндекс здоров'я
sparkutils/quality
A Quality Spark DQ and transformation Library
Scala★ 55 вер. 2026 р.
Apache-2.05 вер. 2026 р. · метрики 2.10.0
PyPI
54Помірнийіндекс здоров'я
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2 527/міс16 серп. 2026 р.
Apache-2.016 серп. 2026 р. · метрики 2.10.0
Go · npm
54Помірнийіндекс здоров'я
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 125 лип. 2026 р.
Apache-2.025 лип. 2026 р. · метрики 2.10.0
Hex
53Помірнийіндекс здоров'я
nshkrdotcom/json_remedy
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Elixir★ 33↓ 2 708/міс17 лип. 2026 р.
MIT17 лип. 2026 р. · метрики 2.10.0
Maven · PyPI
34У зоні ризикуіндекс здоров'я
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 2 82821 лип. 2026 р.
Apache-2.021 лип. 2026 р. · метрики 2.10.0
PyPI
19Критичнийіндекс здоров'я
datafold/data-diff
Compare tables within or across databases
Python★ 2 98813 серп. 2026 р.
MIT13 серп. 2026 р. · метрики 2.10.0