Get all C2C Jobs / hotlists 🔥 Alerts

Senior Iceberg Data Platform Engineer or Apache Iceberg / Lakehouse Operations Engineer

Job Title:  Iceberg DBA / Lakehouse Operations Engineer  

Location : Remote

 

Job Summary

We are seeking a highly skilled Iceberg DBA / Lakehouse Operations Engineer to own the reliability, performance, and operational integrity of the Iceberg data layer powering enterprise analytics and business-critical applications.

This role operates in a large-scale, multi-engine Lakehouse environment, supporting workloads across Spark, Hive, and Impala, and plays a key role in enterprise data modernization initiatives (Hive and Teradata → Iceberg).

The ideal candidate brings deep expertise in Iceberg table operations, metadata management, and query performance optimization, ensuring consistent, high-performance data access across platforms in a cloud-based environment.

This role is critical to ensuring data accuracy and performance—any degradation directly impacts downstream reporting, analytics, and business-critical decision-making.

Key Responsibilities:

Iceberg Data Layer Ownership & Operations

  • Own day-to-day operations of Apache Iceberg tables supporting multiple enterprise applications
  • Ensure data reliability, consistency, and availability across all Lakehouse workloads
  • Maintain operational integrity for datasets at multi-terabyte to petabyte scale

Advanced Table Management & Optimization

  • Execute advanced Iceberg table maintenance and optimization strategies:
    • Compaction (minor/major) and small file mitigation
    • Snapshot expiration and metadata compaction to control metadata growth
    • Orphan file cleanup (vacuum) to maintain storage efficiency
  • Optimize data layout and performance through:
    • File size tuning and distribution strategies
    • Partition evolution and pruning optimization
    • Clustering and ordering techniques (e.g., Z-ordering or similar patterns)

Data Modeling Standards & Lakehouse Design Alignment

  • Support and enforce data modeling best practices aligned with:
    • Normalized data structures (3NF) for source-aligned datasets
    • Medallion architecture (Bronze / Silver / Gold layers) for curated data flows
  • Ensure Iceberg table design aligns with:
    • Data ingestion patterns (raw vs curated layers)
    • Downstream consumption and performance requirements
  • Assist in structuring datasets to balance:
    • Data integrity and normalization
    • Query performance and analytical efficiency
  • Work with data engineering teams to ensure consistent implementation of layered data architecture across multiple applications

Multi-Engine Query Performance & Consistency

  • Ensure consistent and performant query behavior across:
    • Spark (CDE)
    • Hive / Impala (CDW)
  • Troubleshoot and resolve:
    • Query performance bottlenecks
    • Metadata inconsistencies across engines
    • Inefficient execution plans and scan patterns

Metadata & Data Lifecycle Management

  • Manage Iceberg metadata to ensure:
    • Efficient scaling and performance
    • Consistent table state across engines
  • Execute lifecycle operations:
    • Data retention and archival policies
    • Snapshot lifecycle management and cleanup
    • Time-travel optimization and maintenance

Production Support, Incident Resolution & On-Call

  • Provide L2/L3 support for data-related production issues across Iceberg-based Lakehouse workloads
  • Participate in on-call rotation to support critical data platforms and ensure timely response to incidents
  • Respond to and resolve P1/P2 production incidents within defined SLAs, minimizing impact to downstream applications and reporting
  • Troubleshoot:
    • Data inconsistencies and reporting discrepancies
    • Query failures and performance degradation
  • Perform root cause analysis (RCA) and implement preventive measures to avoid recurring issues
  • Collaborate with platform and application teams during incident triage and resolution

             Security & Data Governance Support

  • Support fine-grained access control using:
    • Ranger policies and RBAC
  • Own and ensure data validation, reconciliation, and accuracy between source and Iceberg datasets
  • Ensure secure and compliant access to data across applications

Required Skills

  • Strong hands-on experience with Apache Iceberg and/or Hive-based data lakes
  • Understanding of data modeling concepts (normal forms) and modern Lakehouse patterns (Medallion architecture)
  • Expertise in:
    • Table-level optimization and performance tuning
    • Large-scale data management (TB/PB scale)
  • Experience with:
    • Spark SQL, Hive, Impala, NiFI, Trino
  • Strong understanding of:
    • Partitioning strategies
    • File formats (Parquet/ORC)
    • Distributed query processing



Thanks & Regards,
Maddula Venkateshwara Reddy | ICS Global Soft
Senior. US IT RECRUITER
venkatreddy61996@gmail.com

About Author

I’m Monica Kerry, a passionate SEO and Digital Marketing Specialist with over 9 years of experience helping businesses grow their online presence. From SEO strategy, keyword research, content optimization, and link building to social media marketing and PPC campaigns, I specialize in driving organic traffic, boosting rankings, and increasing conversions. My mission is to empower brands with result-oriented digital marketing solutions that deliver measurable success.

Leave a Reply

Your email address will not be published. Required fields are marked *

×

Post your C2C job instantly

Quick & easy posting in 10 seconds

Keep it concise - you can add details later
Please use your company/professional email address
Simple math question to prevent spam