Todas las etiquetas
Etiqueta del catálogo

#data-pipeline

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

28 registros
Con la etiqueta «data-pipeline»Ordenado por índice de salud
Maven
98Excepcionalíndice de salud
apache/shardingsphere
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
Java★ 20.8K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
Go
98Excepcionalíndice de salud
rudderlabs/rudder-server
Privacy and Security focused Segment-alternative, in Golang and React
Go★ 447728 ago 2026
Licencia propia28 ago 2026 · métricas 2.10.0
PyPI
97Excepcionalíndice de salud
elementary-data/elementary
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
HTML · Python★ 239013 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
Maven · npm
95Excepcionalíndice de salud
airbytehq/airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
Python · Kotlin★ 22K26 ago 2026
Licencia propia26 ago 2026 · métricas 2.10.0
PyPI · Go · npm
94Excepcionalíndice de salud
bruin-data/ingestr
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
Go★ 3778↓ 87.2K/mes17 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
datajuicer/data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Python★ 68183 ago 2026
Apache-2.03 ago 2026 · métricas 2.10.0
Go
92Excelenteíndice de salud
datazip-inc/olake
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
Go★ 143331 ago 2026
Apache-2.031 ago 2026 · métricas 2.10.0
Go
91Excelenteíndice de salud
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Go★ 60417 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
Maven
91Excelenteíndice de salud
debezium/debezium
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
Java★ 13.1K27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
RubyGems · npm
91Excelenteíndice de salud
multiwoven/multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Ruby · TypeScript★ 166317 jul 2026
AGPL-3.017 jul 2026 · métricas 2.10.0
Go · crates.io · npm +1
90Excelenteíndice de salud
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
Rust · C++★ 94817 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
83Excelenteíndice de salud
Glitni/dlt-saga
Config-driven data ingestion and SCD2 historization framework built on dlt
Python★ 63 ago 2026
Apache-2.03 ago 2026 · métricas 2.10.0
Go
81Excelenteíndice de salud
chalk-ai/chalk-go
Go client for Chalk
Go★ 1016 jul 2026
Apache-2.016 jul 2026 · métricas 2.10.0
PyPI
77Buenoíndice de salud
maltzsama/dagster-authkit
Community Auth System for self-hosted Dagster OSS - simple RBAC, Audit-log, and Session Management
Python★ 63↓ 7705/mes22 ago 2026
Apache-2.022 ago 2026 · métricas 2.10.0
PyPI
75Buenoíndice de salud
olirice/flupy
Fluent data pipelines for python and your shell
Python★ 195↓ 1.6M/mes27 jul 2026
Licencia propia27 jul 2026 · métricas 2.10.0
PyPI
71Buenoíndice de salud
russalo/file-observer
Deterministic file observation for pipelines — one read-only pass over a directory emits a reproducible JSON manifest of every file's type, metadata, structure, and provenance.
Python★ 217 jul 2026
Licencia propia17 jul 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
PFund-Software-Ltd/pfeed
Data Engine for AI/Algo Trading: Download/Stream -> Clean -> Store. Supports Data Lakehouse Architecture. Clean Once and Forget.
Python★ 35↓ 1064/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 2.10.0
PyPI
69Buenoíndice de salud
ebarti/jobstreaming
Concurrent, resumable job collection for Python with typed events, durable checkpoints, normalized models, and a DataFrame API.
Python★ 0↓ 2749/mes29 ago 2026
MIT29 ago 2026 · métricas 2.10.0
PyPI · crates.io
59Moderadoíndice de salud
apitap/apitap-lib
Move whole tables between databases fast — Postgres, MySQL, ClickHouse, BigQuery. Rust engine, one-line Python API, bounded memory.
Rust · Python★ 5018 ago 2026
MIT18 ago 2026 · métricas 2.10.0
PyPI · Go · Maven
59Moderadoíndice de salud
liam0205/pineapple
Multi-runtime DAG pipeline engine — declare in Python, execute in Go / Java / Python, decouple with JSON.
Go · C++ · Java★ 11↓ 1621/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 2.10.0
PyPI
54Moderadoíndice de salud
Jackxiaozhiren/datasentry
Local-first AI copilot for data quality: 39 detectors, six-dimension scoring, AI repair with human approval, drift engine, cron scheduling with distributed worker pool (failover + parallel), MCP/REST/CLI/Web UI. Apache-2.0.
Python★ 2↓ 2527/mes16 ago 2026
Apache-2.016 ago 2026 · métricas 2.10.0
Go · npm
54Moderadoíndice de salud
Tnsor-Labs/brokoli
Brokoli — self-hosted data pipeline orchestration
Go · Svelte★ 125 jul 2026
Apache-2.025 jul 2026 · métricas 2.10.0
PyPI · npm
50Moderadoíndice de salud
N283T/pdb-mine-builder
Build a Mine-schema database from PDB data. Load PDBj structural biology data into PostgreSQL with RDKit chemical search support.
Python★ 0↓ 1096/mes15 jul 2026
MIT15 jul 2026 · métricas 2.10.0
PyPI
48Débilíndice de salud
pydoit/doit
CLI task management & automation tool
Python★ 207713 ago 2026
MIT13 ago 2026 · métricas 2.10.0
Go
47Débilíndice de salud
derickschaefer/reserve
A modern command-line interface (CLI) for interfacing with the Federal Reserve Bank of St. Louis FRED API.
Go★ 220 ago 2026
Licencia propia20 ago 2026 · métricas 2.10.0
PyPI
44Débilíndice de salud
mr3od/bgate-unix
High-performance, hardware-aware Layer 0 deduplication engine. Uses 4-tier short-circuit logic (Size → Fringe → Full Hash) with xxHash128 to eliminate exact duplicates before they enter your processing pipeline. Python 3.11+, uv, SQLite.
Python★ 1↓ 32/mes14 ago 2026
MIT14 ago 2026 · métricas 2.10.0
Go · PyPI
41Débilíndice de salud
zoyluoblue/cronova
Lightweight self-hosted workflow scheduler in a single Go binary — an open-source Airflow/Azkaban alternative. Embedded SQLite, DAGs (cron + dependencies + backfill), polyglot tasks (shell/Python/SQL/JAR/HTTP), web console, REST API, and a built-in MCP server for AI agents. Linux & macOS.
Go · JavaScript★ 023 jul 2026
MIT23 jul 2026 · métricas 2.10.0
Maven · PyPI
34En riesgoíndice de salud
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 282821 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0