Search

Data Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

*No C2C solicitation will be allowed*

\n


\n

Pay: $70 per hour

\n

Type: Onsite Wilmington, DE, McLean, VA, Richmond, VA or NYC hybrid

\n

Duration: 4 months

\n


\n

Key Responsibilities:

\n

    \n
  • Design, develop, and maintain scalable data pipelines using Java, Scala, and Apache Spark.
  • \n

  • Build and support multi-stage AWS data pipelines using S3, AWS Glue, Step Functions, and Lambda.
  • \n

  • Develop and maintain event-driven data architectures using Amazon EventBridge and S3 event notifications.
  • \n

  • Design pipeline workflows that appropriately handle event-triggered reruns, retries, failures, and recovery scenarios.
  • \n

  • Configure and manage AWS Step Functions state machines, including scheduling, orchestration, and synchronization of downstream data processing jobs.
  • \n

  • Develop and optimize AWS Glue jobs, including configuration, resource allocation, and scaling for large-volume data workloads.
  • \n

  • Work with modern data formats and technologies such as Apache Iceberg, including troubleshooting platform and version-specific limitations.
  • \n

  • Implement data quality and validation components to ensure data security, integrity, and consistency across pipelines.
  • \n

  • Identify and address schema drift, data changes, and backfill requirements within data pipelines.
  • \n

  • Troubleshoot pipeline failures, performance issues, and data inconsistencies across distributed AWS environments.
  • \n

  • Collaborate with data engineers, architects, and other technical stakeholders to develop reliable and maintainable data solutions.
  • \n

  • Contribute to technical design decisions, documentation, testing, and continuous improvement of the data platform.
  • \n

\n


\n

Qualifications:

\n

    \n
  • Strong hands-on experience with Java and/or Scala.
  • \n

  • Strong experience with Apache Spark and distributed data processing.
  • \n

  • Hands-on experience building AWS data pipelines using:
  • \n

  • Amazon S3
  • \n

  • AWS Glue
  • \n

  • AWS Step Functions
  • \n

  • AWS Lambda
  • \n

  • Experience with event-driven architectures, particularly EventBridge rules and S3 event notifications.
  • \n

  • Understanding of Step Functions state machines, scheduling, orchestration, and downstream job synchronization.
  • \n

  • Experience configuring and scaling AWS Glue jobs.
  • \n

  • Strong understanding of data quality, schema validation, schema drift, and backfill strategies.
  • \n

  • Experience troubleshooting and supporting production data pipelines.
  • \n

\n


\n

Preferred Qualifications

\n

    \n
  • Experience with Apache Iceberg, including Glue/Iceberg integration and troubleshooting.
  • \n

  • Experience designing resilient pipelines with rerun, retry, and failure-recovery mechanisms.
  • \n

  • Experience implementing data validation or security-focused data quality components.
  • \n

  • Experience working with large-scale cloud data platforms and distributed processing environments.
  • \n

  • Strong problem-solving skills and ability to independently research and resolve technical limitations.
  • \n

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...