npm · PyPI97Excellenthealth index
apache/supersetApache Superset is a Data Visualization and Data Exploration Platform
TypeScript · Python · Jupyter Notebook★ 73.8K↓ 1.1M/moJul 15, 2026
PyPI · npm · Go +193Excellenthealth index
apache/airflowApache Airflow - A platform to programmatically author, schedule, and monitor workflows
Python★ 46.2K↓ 28.9M/moJul 21, 2026
PyPI · npm91Excellenthealth index
PrefectHQ/prefectPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
Python · TypeScript★ 23.4K↓ 392.6K/moJul 21, 2026
PyPI · crates.io88Excellenthealth index
eventual-inc/daftHigh-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Rust · Python · Jupyter Notebook★ 5,635↓ 916.1K/moJul 17, 2026
PyPI · npm · Go +187Excellenthealth index
feast-dev/feastThe Open Source Feature Store for AI/ML
Python · Go★ 7,137↓ 615K/moJul 18, 2026
Maven · npm87Excellenthealth index
kestra-io/kestraEvent Driven Orchestration & Scheduling Platform for Mission Critical Applications
Java · TypeScript · Vue★ 27.4KJul 18, 2026
Go87Excellenthealth index
rudderlabs/rudder-serverPrivacy and Security focused Segment-alternative, in Golang and React
Go★ 4,459Jul 18, 2026
PyPI86Excellenthealth index
great-expectations/great_expectationsAlways know what to expect from your data.
Python★ 11.6KJul 16, 2026
PyPI · npm84Goodhealth index
DataRecce/recceThe data-validation toolkit for enhanced dbt (data build tool) PR review
TypeScript · Python★ 463↓ 0/moJul 14, 2026
PyPI · npm83Goodhealth index
dagster-io/dagsterAn orchestration platform for the development, production, and observation of data assets.
Python · TypeScript★ 15.9K↓ 8.6M/moJul 21, 2026
snowflakedb/snowpark-pythonSnowflake Snowpark Python API
Python★ 339↓ 98.6M/moJul 21, 2026
redpanda-data/connectFancy stream processing made operationally mundane
Go★ 8,711Jul 23, 2026
PyPI · npm79Goodhealth index
dbt-labs/dbt-mcpA MCP (Model Context Protocol) server for interacting with dbt.
Python★ 594↓ 98.4K/moJul 17, 2026
cloudquery/cloudqueryData pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
Go★ 6,458↓ 0/moJul 13, 2026
RubyGems · npm78Goodhealth index
multiwoven/multiwoven🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Ruby · TypeScript★ 1,663Jul 17, 2026
ConduitIO/conduitConduit streams data between data stores. Kafka Connect replacement. No JVM required.
Go★ 604Jul 17, 2026
Go · crates.io · npm +174Goodhealth index
estuary/flow🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
Rust · C++★ 948Jul 17, 2026
SuperCowPowers/workbenchWorkbench: An easy to use Python API for creating and deploying AWS SageMaker Models
Python★ 51↓ 10.1K/moJul 18, 2026
npm · PyPI72Goodhealth index
benseverndev-oss/goldenmatchZero-config entity resolution & record linkage. The zero-tuning Fellegi-Sunter path beats hand-tuned Splink head-to-head and scales from a CSV to a verified 100M-row dedupe in 9.2 min. Fuzzy/exact/probabilistic + PPRL + LLM + identity graph. Python + edge-safe TypeScript (WASM), SQL-native in Postgres & DuckDB, MCP/REST + dbt/Airflow.
Python · TypeScript★ 122Jul 15, 2026
npm · PyPI69Moderatehealth index
AltimateAI/altimate-codeOpen-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.
TypeScript★ 751Jul 18, 2026
Go · npm69Moderatehealth index
datazip-inc/olake-uiFrontend & BFF (Backend for frontend) for Olake. This includes the UI code and backend code for storing the configuration of sync and orchestrating it.
TypeScript · Go★ 32Jul 17, 2026
PyPI · crates.io64Moderatehealth index
BirchKwok/ApexBaseA High-performance HTAP embedded database. A reliable work partner.
Rust · Python★ 3↓ 5,330/moJul 17, 2026
exmergo/dexDex is the agent-native analytics engineering toolkit. Point it at your warehouse and your dbt project. It learns the landscape, authors your transformations, and tells you exactly what to fix when the schema drifts. Built for analytics engineers and data engineers who want more out of their coding agent.
Python★ 12Jul 18, 2026
PyPI63Moderatehealth index
russalo/file-observerDeterministic file observation for pipelines — one read-only pass over a directory emits a reproducible JSON manifest of every file's type, metadata, structure, and provenance.
Python★ 2Jul 17, 2026
PyPI62Moderatehealth index
PaulKov/dponeDeclarative ETL framework for YAML-driven data pipelines
Python★ 2↓ 25.1K/moJul 22, 2026
PyPI61Moderatehealth index
PFund-Software-Ltd/pfeedData Engine for AI/Algo Trading: Download/Stream -> Clean -> Store. Supports Data Lakehouse Architecture. Clean Once and Forget.
Python★ 35↓ 1,064/moJul 20, 2026
npm · PyPI61Moderatehealth index
dataengineeringformachinelearning/dataengineeringformachinelearningDEML — High Throughput Event Platform for Real-Time Detection and Response
TypeScript · JavaScript · Python★ 1↓ 537/moJul 21, 2026
crates.io · PyPI57Moderatehealth index
pmgraham/datagruntDatagrunt is a Python library designed to simplify the way you work with CSV, Excel, and PDF files. It provides a streamlined approach to reading, processing, and transforming your data into various formats, making data manipulation efficient and intuitive.
Python★ 12↓ 20.3K/moJul 13, 2026
PyPI56Moderatehealth index
Query-farm/vgi-lint-checkLint the documentation & metadata quality of VGI (Vector Gateway Interface) data workers — descriptions, column comments, tags, and example queries — with a quality score, per-version baselines, and agent-friendly output.
Python★ 0↓ 11.8K/moJul 19, 2026
crates.io · PyPI56Moderatehealth index
mag1cfrog/delta-funnelFast, lightweight Delta Lake to SQL Server loads without Spark or ODBC.
Rust★ 0Jul 16, 2026