
Large-scale distributed data processing using Python and Apache Spark RDDs, DataFrames, and Spark SQL.
Build cloud big data pipelines using Azure HDInsight, Azure Databricks, and Azure Data Lake Analytics.
Distributed event streaming platform, Kafka topics, producers, consumers, Kafka Connect, and KSQL.
Data warehousing on top of Hadoop HDFS using HiveQL, partitioning, bucketing, and ORC/Parquet formats.