Search

AWS Databricks Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Job Summary

\n


\n

We are looking for an experienced AWS Databricks Data Engineer to design, develop, administer, and optimize data platforms built on Databricks and AWS. The ideal candidate will have strong hands-on experience with Databricks, Apache Spark, AWS cloud services, data pipelines, Unity Catalog, security, infrastructure automation, and cost governance.

\n


\n

The candidate will be responsible for building and supporting scalable data engineering solutions while maintaining security, reliability, performance, and operational standards.

\n


\n

Key Responsibilities

\n


\n

Design, develop, and maintain scalable data engineering solutions using Databricks on AWS.

\n

Develop and optimize data pipelines using Apache Spark, PySpark, SQL, and Delta Lake.

\n

Build ETL/ELT pipelines for batch and near-real-time data processing.

\n

Work with AWS services including S3, IAM, Glue, Lambda, CloudWatch, KMS, and VPC.

\n

Administer and support Databricks workspaces, clusters, jobs, notebooks, compute policies, and access controls.

\n

Implement and manage Unity Catalog for centralized data governance, access control, lineage, auditing, and data discovery.

\n

Configure catalogs, schemas, external locations, storage credentials, permissions, and workspace access.

\n

Implement secure integration between Databricks and AWS S3 using appropriate IAM roles and policies.

\n

Automate Databricks and AWS infrastructure using Terraform / Infrastructure as Code.

\n

Monitor and optimize Databricks cluster performance, workloads, and resource utilization.

\n

Implement Databricks job scheduling, monitoring, alerting, and troubleshooting.

\n

Establish and enforce data security, governance, and compliance standards.

\n

Optimize cloud and Databricks costs through cluster policies, autoscaling, job optimization, and resource monitoring.

\n

Support CI/CD implementation for Databricks notebooks, jobs, workflows, and code.

\n

Work with DevOps teams to integrate Databricks deployments with Git, CI/CD pipelines, and Terraform.

\n

Troubleshoot production data pipeline failures, performance issues, and infrastructure problems.

\n

Collaborate with data architects, application teams, DevOps, security, and business stakeholders.

\n


\n

Required Skills

\n


\n

12+ years of experience in Data Engineering / Big Data / Cloud Data Platforms.

\n

Strong hands-on experience with Databricks on AWS.

\n

Strong knowledge of Apache Spark and PySpark.

\n

Advanced SQL skills.

\n

Experience with Delta Lake / Delta tables.

\n

Strong experience with AWS S3 and IAM.

\n

Experience with Databricks Unity Catalog and data governance.

\n

Experience with Databricks administration and workspace management.

\n

Experience with Terraform / Infrastructure as Code.

\n

Experience with Git and CI/CD.

\n

Strong understanding of data pipeline architecture, ETL/ELT, and data warehousing.

\n

Experience with monitoring, troubleshooting, and performance optimization.

\n


\n

Preferred Skills

\n


\n

AWS Glue

\n

AWS Lambda

\n

AWS CloudWatch

\n

AWS KMS

\n

AWS VPC

\n

Databricks Workflows

\n

Databricks Jobs

\n

Databricks SQL

\n

Delta Live Tables / Lakeflow

\n

MLflow

\n

Kafka or other streaming technologies

\n

Python

\n

Scala

\n

Airflow

\n

Azure DevOps / GitHub Actions / Jenkins

\n

Docker / Kubernetes

\n

Data quality and observability tools

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...