Skip to content
View rohankumardubey's full-sized avatar
:octocat:
Focusing
:octocat:
Focusing

Highlights

  • Pro

Organizations

@StreamNest

Block or report rohankumardubey

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rohankumardubey/README.md

Rohan Dubey, Staff Data Engineer & Staff Platform Engineer building distributed systems, real time data platforms, and high performance engines

Portfolio LinkedIn Email Rohan

whoami

I'm Rohan Dubey, Staff Data Engineer & Staff Platform Engineer. I work where distributed systems, real time data, and performance engineering meet. I turn difficult infrastructure problems into platforms that are fast, resilient, and pleasant to operate.

I am especially interested in stream processing semantics, columnar query execution, data intensive systems, and using Rust, Go, Java, and Python to explore the tradeoffs behind them.

I like systems that stay boring under pressure: observable, recoverable, and predictably fast.

Selected engineering

Dremel, dual Rust and C++ columnar SQL engines SpiderOxide, asynchronous Python crawler with Rust acceleration

DataWizz, local first lakehouse and analytics workspace FlowCore, Rust stream processing engine

The problems I enjoy

Area What I care about
Streaming systems Event time, watermarks, state, backpressure, checkpoints, and recovery
Query engines Columnar formats, vectorized execution, memory layout, and benchmarking
Distributed data Partitioning, consistency, fault tolerance, lakehouse architecture, and observability
Platform engineering Kubernetes native workloads, infrastructure as code, automation, and developer experience

Data and systems stack

Layer Technologies
Streaming and messaging Apache Kafka, Apache Flink, Spark Structured Streaming, Apache Pulsar, Kafka Streams, RabbitMQ, NATS
Compute and processing Apache Spark, Apache Flink, Ray, Apache Hadoop, DataFusion
Query and analytics Trino, Presto, ClickHouse, Dremio, Apache Druid, Apache Pinot, DuckDB
Lakehouse and formats Apache Iceberg, Delta Lake, Apache Hudi, Apache Arrow, Parquet, MinIO
Databases and search PostgreSQL, Cassandra, ScyllaDB, CockroachDB, TiDB, MongoDB, Redis, Elasticsearch
Platform and orchestration Kubernetes, Docker, Terraform, Apache Airflow, Argo, Prometheus, Grafana
Cloud data platforms AWS, Snowflake, BigQuery, Databricks

Rust, Go, Java, Python, Scala, C++, Kafka, Kubernetes, Docker, Terraform, AWS, PostgreSQL, Redis, and Grafana

Spark  ·  Flink  ·  Kafka  ·  Arrow  ·  DataFusion  ·  Iceberg  ·  ClickHouse  ·  gRPC

Open source

Open source is a big part of how I learn and work. I enjoy reading real implementations, reproducing ideas in focused projects, reporting issues, and contributing fixes or improvements when I find something useful. Most of my exploration happens around data engines, streaming systems, storage, databases, and Rust infrastructure.

I like the feedback loop that open source creates. Ideas are discussed in public, designs are tested against real workloads, and improvements can benefit an entire engineering community.

Explore Rohan's repositories View Rohan's open source pull requests

GitHub stats

Rohan Dubey's GitHub statistics Rohan Dubey's most used languages

Contribution streak

Rohan Dubey's GitHub contribution streak

Contribution signal

Rohan Dubey's GitHub contribution activity

Open to conversations about distributed systems, data infrastructure, query engines, and performance engineering.

Start a conversation →

Pinned Loading

  1. flink flink Public

    Forked from apache/flink

    Apache Flink

    Java

  2. Dremel Dremel Public

    Two small columnar SQL engines, one in Rust and one in C++, built to answer the same queries over the same bytes

    Rust

  3. spark spark Public

    Forked from apache/spark

    Apache Spark - A unified analytics engine for large-scale data processing

    Scala 2

  4. ClickHouse ClickHouse Public

    Forked from ClickHouse/ClickHouse

    ClickHouse® is a real-time analytics database management system

    C++

  5. tidb tidb Public

    Forked from pingcap/tidb

    TiDB - the open-source, cloud-native, distributed SQL database designed for modern applications.

    Go

  6. HireOs HireOs Public

    AI-powered interview, screening, and hiring intelligence platform for recruiters and companies.

    Python 1