Todas las etiquetas
Etiqueta del catálogo

#data-engineering

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

36 registros
Con la etiqueta «data-engineering»Ordenado por índice de salud
npm · PyPI
97Excelenteíndice de salud
apache/superset
Apache Superset is a Data Visualization and Data Exploration Platform
TypeScript · Python · Jupyter Notebook★ 73.8K↓ 1.1M/mes15 jul 2026
Apache-2.015 jul 2026 · métricas 1.13.0
PyPI · npm · Go +1
93Excelenteíndice de salud
apache/airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Python★ 46.2K↓ 28.9M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI · npm
91Excelenteíndice de salud
PrefectHQ/prefect
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
Python · TypeScript★ 23.4K↓ 392.6K/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI · crates.io
88Excelenteíndice de salud
eventual-inc/daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Rust · Python · Jupyter Notebook★ 5635↓ 916.1K/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI · npm · Go +1
87Excelenteíndice de salud
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go★ 7137↓ 615K/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
Maven · npm
87Excelenteíndice de salud
kestra-io/kestra
Event Driven Orchestration & Scheduling Platform for Mission Critical Applications
Java · TypeScript · Vue★ 27.4K18 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
Go
87Excelenteíndice de salud
rudderlabs/rudder-server
Privacy and Security focused Segment-alternative, in Golang and React
Go★ 445918 jul 2026
Licencia propia18 jul 2026 · métricas 1.13.0
PyPI · npm
84Buenoíndice de salud
DataRecce/recce
The data-validation toolkit for enhanced dbt (data build tool) PR review
TypeScript · Python★ 463↓ 0/mes14 jul 2026
Apache-2.014 jul 2026 · métricas 1.13.0
PyPI · npm
83Buenoíndice de salud
dagster-io/dagster
An orchestration platform for the development, production, and observation of data assets.
Python · TypeScript★ 15.9K↓ 8.6M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
PyPI
83Buenoíndice de salud
snowflakedb/snowpark-python
Snowflake Snowpark Python API
Python★ 339↓ 98.6M/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 1.13.0
Go
80Buenoíndice de salud
redpanda-data/connect
Fancy stream processing made operationally mundane
Go★ 871123 jul 2026
Licencia propia23 jul 2026 · métricas 1.13.0
PyPI · npm
79Buenoíndice de salud
dbt-labs/dbt-mcp
A MCP (Model Context Protocol) server for interacting with dbt.
Python★ 594↓ 98.4K/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go
78Buenoíndice de salud
cloudquery/cloudquery
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
Go★ 6458↓ 0/mes13 jul 2026
MPL-2.013 jul 2026 · métricas 1.13.0
RubyGems · npm
78Buenoíndice de salud
multiwoven/multiwoven
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Ruby · TypeScript★ 166317 jul 2026
AGPL-3.017 jul 2026 · métricas 1.13.0
Go
76Buenoíndice de salud
ConduitIO/conduit
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Go★ 60417 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
Go · crates.io · npm +1
74Buenoíndice de salud
estuary/flow
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
Rust · C++★ 94817 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI
72Buenoíndice de salud
SuperCowPowers/workbench
Workbench: An easy to use Python API for creating and deploying AWS SageMaker Models
Python★ 51↓ 10.1K/mes18 jul 2026
MIT18 jul 2026 · métricas 1.13.0
npm · PyPI
72Buenoíndice de salud
benseverndev-oss/goldenmatch
Zero-config entity resolution & record linkage. The zero-tuning Fellegi-Sunter path beats hand-tuned Splink head-to-head and scales from a CSV to a verified 100M-row dedupe in 9.2 min. Fuzzy/exact/probabilistic + PPRL + LLM + identity graph. Python + edge-safe TypeScript (WASM), SQL-native in Postgres & DuckDB, MCP/REST + dbt/Airflow.
Python · TypeScript★ 12215 jul 2026
MIT15 jul 2026 · métricas 1.13.0
npm · PyPI
69Moderadoíndice de salud
AltimateAI/altimate-code
Open-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.
TypeScript★ 75118 jul 2026
MIT18 jul 2026 · métricas 1.13.0
Go · npm
69Moderadoíndice de salud
datazip-inc/olake-ui
Frontend & BFF (Backend for frontend) for Olake. This includes the UI code and backend code for storing the configuration of sync and orchestrating it.
TypeScript · Go★ 3217 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
PyPI · crates.io
64Moderadoíndice de salud
BirchKwok/ApexBase
A High-performance HTAP embedded database. A reliable work partner.
Rust · Python★ 3↓ 5330/mes17 jul 2026
Apache-2.017 jul 2026 · métricas 1.13.0
63Moderadoíndice de salud
exmergo/dex
Dex is the agent-native analytics engineering toolkit. Point it at your warehouse and your dbt project. It learns the landscape, authors your transformations, and tells you exactly what to fix when the schema drifts. Built for analytics engineers and data engineers who want more out of their coding agent.
Python★ 1218 jul 2026
Apache-2.018 jul 2026 · métricas 1.13.0
PyPI
63Moderadoíndice de salud
russalo/file-observer
Deterministic file observation for pipelines — one read-only pass over a directory emits a reproducible JSON manifest of every file's type, metadata, structure, and provenance.
Python★ 217 jul 2026
Licencia propia17 jul 2026 · métricas 1.13.0
PyPI
62Moderadoíndice de salud
PaulKov/dpone
Declarative ETL framework for YAML-driven data pipelines
Python★ 2↓ 25.1K/mes22 jul 2026
Apache-2.022 jul 2026 · métricas 1.13.0
PyPI
61Moderadoíndice de salud
PFund-Software-Ltd/pfeed
Data Engine for AI/Algo Trading: Download/Stream -> Clean -> Store. Supports Data Lakehouse Architecture. Clean Once and Forget.
Python★ 35↓ 1064/mes20 jul 2026
Apache-2.020 jul 2026 · métricas 1.13.0
crates.io · PyPI
57Moderadoíndice de salud
pmgraham/datagrunt
Datagrunt is a Python library designed to simplify the way you work with CSV, Excel, and PDF files. It provides a streamlined approach to reading, processing, and transforming your data into various formats, making data manipulation efficient and intuitive.
Python★ 12↓ 20.3K/mes13 jul 2026
MIT13 jul 2026 · métricas 1.13.0
PyPI
56Moderadoíndice de salud
Query-farm/vgi-lint-check
Lint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/mes19 jul 2026
Licencia propia19 jul 2026 · métricas 1.13.0
crates.io · PyPI
56Moderadoíndice de salud
mag1cfrog/delta-funnel
Fast, lightweight Delta Lake to SQL Server loads without Spark or ODBC.
Rust★ 016 jul 2026
Apache-2.016 jul 2026 · métricas 1.13.0