
Large-scale distributed data processing using Python and Apache Spark RDDs, DataFrames, and Spark SQL.
Distributed event streaming platform, Kafka topics, producers, consumers, Kafka Connect, and KSQL.
Install, configure, monitor, and maintain Hadoop HDFS and YARN clusters using Ambari/Cloudera.
Data warehousing on top of Hadoop HDFS using HiveQL, partitioning, bucketing, and ORC/Parquet formats.