• Spark hands on with optimization and troubleshooting
• Proficiency with the AWS data ecosystem (EMR, S3, Glue, Lambda, Redshift, etc.).
• Ability to write and optimize complex SQL queries
• Well versed with either Python or Scala
• Ability to develop and manage data pipelines
• good understanding of data warehousing and data lake concept
• Monitoring services and capabilities