Todas las etiquetas
Etiqueta del catálogo

#data-pipeline

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

14 registros
Con la etiqueta «data-pipeline»Ordenado por índice de salud
Go
87Excelenteíndice de salud
rudderlabs/rudder-server
Privacy and Security focused Segment-alternative, in Golang and React
Go★ 445918 jul 2026
Licencia propia18 jul 2026 · métricas 1.13.0
Maven
86Excelenteíndice de salud
apache/shardingsphere
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
Java★ 20.8K20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
PyPI · Go · npm
79Buenoíndice de salud
bruin-data/ingestr
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
Go★ 3778↓ 87.2K/mes17 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
RubyGems · npm
78Buenoíndice de salud
multiwoven/multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Ruby · TypeScript★ 166317 jul 2026
AGPL-3.017 jul 2026 · métricas 1.13.0
Maven
77Buenoíndice de salud
debezium/debezium
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
Java★ 12.9K20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
Go
76Buenoíndice de salud
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Go★ 60417 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go · crates.io · npm +1
74Buenoíndice de salud
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
Rust · C++★ 94817 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
Go
69Moderadoíndice de salud
chalk-ai/chalk-go
Go client for Chalk
Go★ 1016 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
russalo/file-observer
Deterministic file observation for pipelines — one read-only pass over a directory emits a reproducible JSON manifest of every file's type, metadata, structure, and provenance.
Python★ 217 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI
61Moderadoíndice de salud
PFund-Software-Ltd/pfeed
Data Engine for AI/Algo Trading: Download/Stream -> Clean -> Store. Supports Data Lakehouse Architecture. Clean Once and Forget.
Python★ 35↓ 1064/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
PyPI · Go · Maven
56Moderadoíndice de salud
liam0205/pineapple
Multi-runtime DAG pipeline engine — declare in Python, execute in Go / Java / Python, decouple with JSON.
Go · C++ · Java★ 11↓ 1621/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI · npm
50Moderadoíndice de salud
N283T/pdb-mine-builder
Build a Mine-schema database from PDB data. Load PDBj structural biology data into PostgreSQL with RDKit chemical search support.
Python★ 0↓ 1096/mes15 jul 2026
MIT15 jul 2026 · métricas 1.13.0
PyPI
50Moderadoíndice de salud
maltzsama/dagster-authkit
Community Auth System for self-hosted Dagster OSS - simple RBAC, Audit-log, and Session Management
Python★ 55↓ 7705/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
Maven · PyPI
38En riesgoíndice de salud
whylabs/whylogs
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
Jupyter Notebook · Python · HTML★ 282821 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0