全部标签
目录标签

#data-pipeline

公开记录中带有此标签的全部仓库——标签来自其 GitHub 主题或软件包注册表发布的关键词。健康度量遵循与记录其余部分相同的版本化方法论。

28 条记录
标签为“data-pipeline”按健康指数排序
Maven
98卓越健康指数
apache/shardingsphere
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
Java★ 20.8K2026年8月5日
Apache-2.02026年8月5日 · 指标 2.10.0
Go
98卓越健康指数
rudderlabs/rudder-server
Privacy and Security focused Segment-alternative, in Golang and React
Go★ 4,4772026年8月28日
自定义许可证2026年8月28日 · 指标 2.10.0
PyPI
97卓越健康指数
elementary-data/elementary
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
HTML · Python★ 2,3902026年8月13日
Apache-2.02026年8月13日 · 指标 2.10.0
Maven · npm
95卓越健康指数
airbytehq/airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
Python · Kotlin★ 22K2026年8月26日
自定义许可证2026年8月26日 · 指标 2.10.0
PyPI · Go · npm
94卓越健康指数
bruin-data/ingestr
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
Go★ 3,778↓ 87.2K/月2026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
PyPI
94卓越健康指数
datajuicer/data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Python★ 6,8182026年8月3日
Apache-2.02026年8月3日 · 指标 2.10.0
Go
92优秀健康指数
datazip-inc/olake
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
Go★ 1,4332026年8月31日
Apache-2.02026年8月31日 · 指标 2.10.0
Go
91优秀健康指数
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Go★ 6042026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
Maven
91优秀健康指数
debezium/debezium
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
Java★ 13.1K2026年8月27日
Apache-2.02026年8月27日 · 指标 2.10.0
RubyGems · npm
91优秀健康指数
multiwoven/multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Ruby · TypeScript★ 1,6632026年7月17日
AGPL-3.02026年7月17日 · 指标 2.10.0
Go · crates.io · npm +1
90优秀健康指数
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
Rust · C++★ 9482026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
PyPI
83优秀健康指数
Glitni/dlt-saga
Config-driven data ingestion and SCD2 historization framework built on dlt
Python★ 62026年8月3日
Apache-2.02026年8月3日 · 指标 2.10.0
Go
81优秀健康指数
chalk-ai/chalk-go
Go client for Chalk
Go★ 102026年7月16日
Apache-2.02026年7月16日 · 指标 2.10.0
PyPI
77良好健康指数
maltzsama/dagster-authkit
Community Auth System for self-hosted Dagster OSS - simple RBAC, Audit-log, and Session Management
Python★ 63↓ 7,705/月2026年8月22日
Apache-2.02026年8月22日 · 指标 2.10.0
PyPI
75良好健康指数
olirice/flupy
Fluent data pipelines for python and your shell
Python★ 195↓ 1.6M/月2026年7月27日
自定义许可证2026年7月27日 · 指标 2.10.0
PyPI
71良好健康指数
russalo/file-observer
Deterministic file observation for pipelines — one read-only pass over a directory emits a reproducible JSON manifest of every file's type, metadata, structure, and provenance.
Python★ 22026年7月17日
自定义许可证2026年7月17日 · 指标 2.10.0
PyPI
69良好健康指数
PFund-Software-Ltd/pfeed
Data Engine for AI/Algo Trading: Download/Stream -> Clean -> Store. Supports Data Lakehouse Architecture. Clean Once and Forget.
Python★ 35↓ 1,064/月2026年7月20日
Apache-2.02026年7月20日 · 指标 2.10.0
PyPI
69良好健康指数
ebarti/jobstreaming
Concurrent, resumable job collection for Python with typed events, durable checkpoints, normalized models, and a DataFrame API.
Python★ 0↓ 2,749/月2026年8月29日
MIT2026年8月29日 · 指标 2.10.0
PyPI · crates.io
59中等健康指数
apitap/apitap-lib
Move whole tables between databases fast — Postgres, MySQL, ClickHouse, BigQuery. Rust engine, one-line Python API, bounded memory.
Rust · Python★ 502026年8月18日
MIT2026年8月18日 · 指标 2.10.0
PyPI · Go · Maven
59中等健康指数
liam0205/pineapple
Multi-runtime DAG pipeline engine — declare in Python, execute in Go / Java / Python, decouple with JSON.
Go · C++ · Java★ 11↓ 1,621/月2026年7月17日
Apache-2.02026年7月17日 · 指标 2.10.0
PyPI
54中等健康指数
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2,527/月2026年8月16日
Apache-2.02026年8月16日 · 指标 2.10.0
Go · npm
54中等健康指数
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 12026年7月25日
Apache-2.02026年7月25日 · 指标 2.10.0
PyPI · npm
50中等健康指数
N283T/pdb-mine-builder
Build a Mine-schema database from PDB data. Load PDBj structural biology data into PostgreSQL with RDKit chemical search support.
Python★ 0↓ 1,096/月2026年7月15日
MIT2026年7月15日 · 指标 2.10.0
PyPI
48薄弱健康指数
pydoit/doit
CLI task management & automation tool
Python★ 2,0772026年8月13日
MIT2026年8月13日 · 指标 2.10.0
Go
47薄弱健康指数
derickschaefer/reserve
A modern command-line interface (CLI) for interfacing with the Federal Reserve Bank of St. Louis FRED API.
Go★ 22026年8月20日
自定义许可证2026年8月20日 · 指标 2.10.0
PyPI
44薄弱健康指数
mr3od/bgate-unix
High-performance, hardware-aware Layer 0 deduplication engine. Uses 4-tier short-circuit logic (Size → Fringe → Full Hash) with xxHash128 to eliminate exact duplicates before they enter your processing pipeline. Python 3.11+, uv, SQLite.
Python★ 1↓ 32/月2026年8月14日
MIT2026年8月14日 · 指标 2.10.0
Go · PyPI
41薄弱健康指数
zoyluoblue/cronova
Lightweight self-hosted workflow scheduler in a single Go binary — an open-source Airflow/Azkaban alternative. Embedded SQLite, DAGs (cron + dependencies + backfill), polyglot tasks (shell/Python/SQL/JAR/HTTP), web console, REST API, and a built-in MCP server for AI agents. Linux & macOS.
Go · JavaScript★ 02026年7月23日
MIT2026年7月23日 · 指标 2.10.0
Maven · PyPI
34存在风险健康指数
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 2,8282026年7月21日
Apache-2.02026年7月21日 · 指标 2.10.0