Todas las etiquetas
Etiqueta del catálogo

#big-data

Todos los repositorios del registro público que llevan esta etiqueta, procedente de sus topics de GitHub o de las palabras clave que publican sus registros de paquetes. La salud se mide con la misma metodología versionada que el resto del registro.

39 registros
Con la etiqueta «big-data»Ordenado por índice de salud
crates.io · PyPI
99Excepcionalíndice de salud
apache/datafusion
Apache DataFusion SQL Query Engine
Rust★ 9211↓ 18.6K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI · crates.io
98Excepcionalíndice de salud
Eventual-Inc/Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Rust · Python · Jupyter Notebook★ 5731↓ 6.3M/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI · npm
98Excepcionalíndice de salud
delta-io/delta
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
Scala · Java★ 8926↓ 37M/mes2 ago 2026
Apache-2.02 ago 2026 · métricas 2.10.0
PyPI · npm · Go +1
98Excepcionalíndice de salud
feast-dev/feast
The Open Source Feature Store for AI/ML
Python · Go · TypeScript★ 7233↓ 841.1K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
Maven
98Excepcionalíndice de salud
trinodb/trino
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Java★ 13.2K27 ago 2026
Apache-2.027 ago 2026 · métricas 2.10.0
Go · Maven
97Excepcionalíndice de salud
apache/beam
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python★ 865328 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
Maven · PyPI
97Excepcionalíndice de salud
prestodb/presto
The official home of the Presto distributed SQL query engine for big data
Java★ 16.7K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
Maven · RubyGems
97Excepcionalíndice de salud
vespa-engine/vespa
The AI search platform
Java · C++★ 707228 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
Maven · PyPI
96Excepcionalíndice de salud
StarRocks/starrocks
The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.
Java · C++★ 12K12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
PyPI · Maven · npm
96Excepcionalíndice de salud
apache/tsfile
Apache TsFile
Java · C++★ 19527 jul 2026
Apache-2.027 jul 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
graphframes/graphframes
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
Scala · Python★ 120113 ago 2026
Apache-2.013 ago 2026 · métricas 2.10.0
PyPI
96Excepcionalíndice de salud
man-group/arcticdb
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
C++ · Python★ 242416 jul 2026
Licencia propia16 jul 2026 · métricas 2.10.0
Maven · npm · PyPI +1
95Excepcionalíndice de salud
apache/spark
Apache Spark - A unified analytics engine for large-scale data processing
Scala · Python★ 43.8K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
npm · PyPI
95Excepcionalíndice de salud
microsoft/SynapseML
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
Scala · Python★ 523612 ago 2026
MIT12 ago 2026 · métricas 2.10.0
crates.io · PyPI
94Excepcionalíndice de salud
apache/datafusion-ballista
Apache DataFusion Ballista Distributed Query Engine
Rust★ 2093↓ 58/mes21 jul 2026
Apache-2.021 jul 2026 · métricas 2.10.0
PyPI
94Excepcionalíndice de salud
cython/cython
The most widely used Python to C compiler
Cython · Python★ 10.8K28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
Maven
93Excepcionalíndice de salud
alibaba/fastjson2
🚄 FASTJSON2 is a Java JSON library with excellent performance.
Java★ 439728 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI
93Excepcionalíndice de salud
delta-io/delta-sharing
An open protocol for secure data sharing
Scala · Python★ 956↓ 1.4M/mes21 ago 2026
Apache-2.021 ago 2026 · métricas 2.10.0
Maven
93Excepcionalíndice de salud
hazelcast/hazelcast
Hazelcast is a unified real-time data platform combining stream processing with a fast data store, allowing customers to act instantly on data-in-motion for real-time insights.
Java★ 660122 ago 2026
Licencia propia22 ago 2026 · métricas 2.10.0
crates.io
93Excepcionalíndice de salud
reductstore/reductstore
High-performance, time-indexed object storage for robotics and industrial IoT
Rust★ 368↓ 192/mes30 ago 2026
Apache-2.030 ago 2026 · métricas 2.10.0
Maven · PyPI
91Excelenteíndice de salud
apache/flink
Apache Flink
Java★ 26.2K5 ago 2026
Apache-2.05 ago 2026 · métricas 2.10.0
Maven · RubyGems
91Excelenteíndice de salud
apache/orc
Apache ORC - the smallest, fastest columnar storage for Hadoop workloads
Java · C++★ 76822 jul 2026
Apache-2.022 jul 2026 · métricas 2.10.0
npm
91Excelenteíndice de salud
catboost/catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
C++ · Python★ 9064↓ 7903/mes12 ago 2026
Apache-2.012 ago 2026 · métricas 2.10.0
Maven · RubyGems
90Excelenteíndice de salud
apache/calcite
Apache Calcite
Java★ 517221 ago 2026
Apache-2.021 ago 2026 · métricas 2.10.0
Go
90Excelenteíndice de salud
ytsaurus/ytsaurus-k8s-operator
Kubernetes operator for YTsaurus.
Go★ 469 ago 2026
Licencia propia9 ago 2026 · métricas 2.10.0
npm
88Excelenteíndice de salud
hazelcast/hazelcast-nodejs-client
Hazelcast Node.js Client
TypeScript · JavaScript★ 152↓ 25.5K/mes28 ago 2026
Apache-2.028 ago 2026 · métricas 2.10.0
PyPI · Maven · npm +1
87Excelenteíndice de salud
h2oai/h2o-3
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.
Jupyter Notebook · HTML · Java★ 7495↓ 3086/mes24 ago 2026
Apache-2.024 ago 2026 · métricas 2.10.0
PyPI
86Excelenteíndice de salud
SuperCowPowers/workbench
Workbench: An easy to use Python API for creating and deploying AWS SageMaker Models
Python★ 51↓ 10.1K/mes18 jul 2026
MIT18 jul 2026 · métricas 2.10.0
PyPI · crates.io
86Excelenteíndice de salud
mabel-dev/opteryx
🦖 A SQL-on-everything Query Engine you can execute over multiple databases and file formats. Query your data, where it lives.
Python★ 113↓ 11.3K/mes18 jul 2026
Apache-2.018 jul 2026 · métricas 2.10.0
crates.io · Maven · PyPI
80Excelenteíndice de salud
apache/paimon-vector-index
Apache Paimon Vector Index: pure Rust IVF-PQ for data lake vector search.
Rust★ 20↓ 3014/mes30 jul 2026
Apache-2.030 jul 2026 · métricas 2.10.0