
Large-scale distributed data processing using Python and Apache Spark RDDs, DataFrames, and Spark SQL.
Build cloud big data pipelines using Azure HDInsight, Azure Databricks, and Azure Data Lake Analytics.
Install, configure, monitor, and maintain Hadoop HDFS and YARN clusters using Ambari/Cloudera.
Data warehousing on top of Hadoop HDFS using HiveQL, partitioning, bucketing, and ORC/Parquet formats.