Search

Lead Data Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Job Description

We're seeking a senior hands-on Lead Data Engineer to own the end-to-end technical design of data and AI layers for a centralized, AI-first enterprise Data Hub built on Azure Databricks. The role will lead lakehouse architecture, metadata-driven ingestion, AI-augmented data engineering, data quality, governance, and production engineering while providing technical leadership and mentoring to the engineering team.

Job Description:

Candidates who have not written or reviewed production code in the past year are not a fit. WHAT THE ROLE OWNS - Data platform architecture and engineering: lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, retention). - Metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka, Spark Structured Streaming). - Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning and cost guardrails. - CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; observability with Azure Monitor and Log Analytics. - AI-augmented ingestion and canonical mapping: auto-generated bridge documents, DML, canonical table definitions; AI-assisted source-to-canonical mapping with human review gate. - AI-driven data quality, anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data. - Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security. - Governance and leadership: Unity Catalog (lineage, access control, PII standards), Architecture Review Boards and AI governance forums, mentoring engineers, documentation. MUST-HAVE SKILLS & EXPERIENCE - Expert-level Python, Scala and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimizing large-scale batch and streaming pipelines using Delta Lake. - Strong SQL and data modelling (dimensional and normalised), schema design, data contracts. - Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving. - Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks), Azure Event Hubs. - 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering. - Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback. - Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving). - Metadata-driven frameworks: schema inference, data profiling, lineage, catalogs. - 12-18 years of total experience in data engineering / data platform delivery. - Proven enterprise-scale delivery of a medallion / lakehouse architecture. - Azure security and governance: Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling. - CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines. - Clear technical writing and ability to present and defend designs to engineers and non-technical stakeholders. STRONGLY PREFERRED - Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling over a lakehouse). - Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale. - ML-based anomaly detection on time-series or transactional financial data. - Financial services or insurance domain (finance close, GL, subledger, reconciliation, actuarial data). - LLMOps / MLOps: model and prompt versioning, cost governance, observability. - Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification. - dbt, Great Expectations or similar; Workday, Prism or Accounting Center exposure.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...