PySpark, Azure Big Data, Apache Kafka, Hadoop Administration, Hive

Large-scale distributed data processing using Python and Apache Spark RDDs, DataFrames, and Spark SQL.

Build cloud big data pipelines using Azure HDInsight, Azure Databricks, and Azure Data Lake Analytics.

Distributed event streaming platform, Kafka topics, producers, consumers, Kafka Connect, and KSQL.

Install, configure, monitor, and maintain Hadoop HDFS and YARN clusters using Ambari/Cloudera.

Data warehousing on top of Hadoop HDFS using HiveQL, partitioning, bucketing, and ORC/Parquet formats.

Distributed NoSQL column-oriented database running on top of HDFS for real-time read/write access.

Distributed real-time computation system, Spouts, Bolts, Topologies, and fault-tolerant streaming.

High-level data flow platform and Pig Latin scripting language for transforming large datasets on Hadoop.