Title: Lead Azure Data Engineer
Location: Remote
Databricks certification must.
Databricks Engineer to provide L2/L3 production support for an enterprise data platform built on Azure Databricks and Delta Lake. The role will be responsible for monitoring, troubleshooting, incident resolution, data pipeline support, performance optimization, and ensuring the reliability and availability of critical data workloads.
The engineer will work closely with L1 support, Data Engineers, Data Architects, BI/Reporting teams, application teams, infrastructure teams, and business stakeholders to resolve production issues and maintain platform stability.
Key Responsibilities
1. Production Support & Incident Management
- Provide L2/L3 technical support for Azure Databricks data platform production workloads.
- Analyze and resolve production incidents related to data pipelines, notebooks, jobs, clusters, Delta tables, and data processing.
- Perform root-cause analysis for recurring and high-severity incidents.
- Support incident triage, prioritization, escalation, and resolution within defined SLAs.
- Participate in major incident management and provide technical inputs for service restoration.
- Maintain incident resolution documentation and contribute to the Known Error Database (KEDB).
2. Databricks & Data Pipeline Support
- Monitor and troubleshoot Databricks Jobs, Workflows, notebooks, clusters, and data pipelines.
- Troubleshoot failures in ETL/ELT pipelines and identify issues across source, transformation, and target layers.
- Analyze Spark job failures, executor issues, memory constraints, timeouts, and performance bottlenecks.
- Support Delta Lake tables, including data quality, schema evolution, partitioning, and optimization.
- Perform data validation and reconciliation following production incidents or reruns.
- Manage production job reruns and recovery activities following approved operational procedures.
3. Performance & Platform Optimization
- Analyze Databricks workload performance and identify opportunities for optimization.
- Optimize Spark jobs, SQL queries, cluster configurations, partitioning, and data processing strategies.
- Troubleshoot inefficient data processing and excessive compute utilization.
- Support Delta table optimization activities such as compaction and maintenance.
- Identify recurring performance issues and recommend permanent fixes.
4. Monitoring & Proactive Support
- Monitor production data pipelines, workflows, and critical data processes.
- Analyze job failures, alerts, logs, and operational metrics.
- Identify potential production issues before they impact business users.
- Develop and maintain operational dashboards and monitoring mechanisms where required.
- Support proactive health checks for critical data platform components.
5. Change & Release Support
- Support production deployments and changes related to Databricks workloads.
- Review technical implementation plans, deployment procedures, and rollback plans.
- Validate changes following production deployment.
- Support emergency changes when required.
- Ensure appropriate documentation and approvals are available before production implementation.
6. Data Quality & Reconciliation
- Investigate data discrepancies and pipeline-related data quality issues.
- Perform source-to-target validation and reconciliation.
- Identify data loss, duplication, transformation, or latency issues.
- Collaborate with Data Engineering and Business teams to resolve data quality issues.
- Implement or recommend automated data validation checks where appropriate.
7. Problem Management & Continuous Improvement
- Identify recurring incidents and initiate problem management activities.
- Conduct Root Cause Analysis (RCA) and define corrective/preventive actions.
- Identify opportunities to automate repetitive support activities.
- Develop reusable scripts, utilities, and operational tools to improve support efficiency.
- Contribute to reduction of recurring incidents and manual operational effort.
8. Documentation & Knowledge Management
- Maintain operational runbooks, SOPs, troubleshooting guides, and support documentation.
- Document known errors, workarounds, and permanent resolutions.
- Maintain application/data pipeline dependency information.
- Support knowledge transfer between L1, L2, L3, development, and architecture teams.
Maddula Venkateshwara Reddy | ICS Global Soft
Senior. US IT RECRUITER
venkatreddy61996@gmail.com
—