Databricks Engineer / Databricks Architect
Remote
Long Term
Contract
Databricks Architect leads the data-layer design and delivery for enterprise AI use cases, serving as the technical authority on the Databricks platform. This role owns the end-to-end architecture of the data foundation that powers AI/BI solutions — from data discovery and readiness through governed, secure delivery — ensuring every solution aligns with enterprise standards. The Architect explicitly owns security design for the platform, including the Unity Catalog governance model, fine-grained access controls, and PII handling for natural-language query experiences.
Key Responsibilities
AI Use Case Data-Layer Design & Delivery
• Lead the architecture, design, and delivery of the data layer supporting prioritized AI use cases, from source ingestion through curated, consumption-ready data products.
• Translate AI/BI use case requirements into concrete data models, pipelines, and platform patterns on Databricks (Delta Lake, medallion architecture, streaming and batch ingestion).
• Define and enforce data quality, lineage, and observability standards so downstream AI models and natural-language query experiences operate on trusted data.
• Establish reusable reference architectures and design patterns that accelerate delivery across successive use cases.
Data Discovery & Readiness
• Co-facilitate live data discovery sessions with the Lead AI/BI Engineer, working directly with business stakeholders and data owners to identify, profile, and validate candidate data sources.
• Assess data readiness for each use case — completeness, quality, granularity, latency, and access — and produce clear readiness findings with remediation plans.
• Drive resolution of data readiness issues, coordinating with source-system owners, data engineering teams, and governance stakeholders to close gaps on schedule.
• Maintain a data discovery playbook and artifacts (source inventories, profiling results, gap logs) that make each engagement faster than the last.
Databricks Platform Architecture
• Architect the Databricks platform solution — workspace topology, compute strategy, storage layout, networking, and CI/CD — aligned to enterprise standards.
• Define environment strategy (dev/test/prod), promotion paths, and infrastructure-as-code practices for repeatable, auditable deployments.
• Advise on cost optimization, performance tuning, and capacity planning across clusters, SQL warehouses, and serverless compute.
• Stay current on the Databricks roadmap (Unity Catalog, Genie/AI-BI, Delta Sharing, serverless) and guide adoption decisions.
Security Design (Explicit Ownership)
• Own the platform security design end to end, including identity integration, workspace access, secrets management, and network isolation.
• Design and implement the Unity Catalog governance model: catalog/schema structure, ownership model, access policies, tagging, and lineage.
• Define and implement row-level and column-level security, dynamic data masking, and attribute-based access controls to enforce least-privilege data access.
• Own PII handling for natural-language query experiences — classification, masking/tokenization strategies, and guardrails that prevent sensitive data exposure through conversational and generative interfaces.
• Partner with enterprise security, privacy, and compliance teams to ensure designs satisfy regulatory and audit requirements, and document controls for review.
Collaboration & Leadership
• Serve as the primary data-architecture counterpart to the Lead AI/BI Engineer, aligning the data layer with semantic models and AI/BI experiences.
• Provide technical direction and design review for data engineers delivering pipelines and data products.
• Communicate architecture decisions, trade-offs, and risks clearly to both technical teams and business stakeholders.
Required Qualifications
• 8+ years in data architecture or data engineering, with 3+ years architecting solutions on Databricks in production environments.
• Deep expertise with the Databricks Lakehouse platform: Delta Lake, Unity Catalog, Databricks SQL, workflows/jobs, and medallion architectures.
• Demonstrated ownership of data security and governance design, including Unity Catalog governance models, row/column-level security, and data masking.
• Hands-on experience with PII classification and protection strategies, ideally in the context of AI, natural-language query, or conversational analytics workloads.
• Strong data modeling skills (dimensional, data vault, or domain-driven data product design) and proficiency in SQL and Python (PySpark).
• Experience with at least one major cloud platform (Azure, AWS, or GCP), including networking, identity (e.g., Entra ID/IAM), and storage services.
• Proven ability to facilitate discovery workshops and translate ambiguous business needs into actionable data designs.
• Experience delivering within enterprise architecture and governance frameworks and standards.
Preferred Qualifications
• Databricks certifications (e.g., Data Engineer Professional, Platform Architect accreditation).
• Experience supporting GenAI/LLM or AI-BI (e.g., Databricks Genie) use cases, including retrieval patterns and semantic layer design.
• Familiarity with data privacy regulations (GDPR, CCPA, HIPAA as applicable) and audit/compliance processes.
• Experience with infrastructure-as-code (Terraform) and CI/CD for Databricks (Asset Bundles, GitHub Actions/Azure DevOps).
• Background in consulting or multi-stakeholder delivery environments.
Success Measures
• AI use case data layers delivered on schedule with documented, standards-aligned architectures.
• Data readiness issues identified early and resolved without derailing delivery timelines.
• Zero PII exposure incidents through natural-language query or AI interfaces; security designs passing enterprise security and audit review.
• Unity Catalog governance model adopted as the enterprise pattern, with measurable reuse across use cases.
Munesh
770-838-3829,
CYBER SPHERE LLC