Job Description
*No C2C solicitation will be allowed*
\n
\n
Pay: $70 per hour
\n
Type: Onsite Wilmington, DE, McLean, VA, Richmond, VA or NYC hybrid
\n
Duration: 4 months
\n
\n
Key Responsibilities:
\n
- \n
- Design, develop, and maintain scalable data pipelines using Java, Scala, and Apache Spark.
- Build and support multi-stage AWS data pipelines using S3, AWS Glue, Step Functions, and Lambda.
- Develop and maintain event-driven data architectures using Amazon EventBridge and S3 event notifications.
- Design pipeline workflows that appropriately handle event-triggered reruns, retries, failures, and recovery scenarios.
- Configure and manage AWS Step Functions state machines, including scheduling, orchestration, and synchronization of downstream data processing jobs.
- Develop and optimize AWS Glue jobs, including configuration, resource allocation, and scaling for large-volume data workloads.
- Work with modern data formats and technologies such as Apache Iceberg, including troubleshooting platform and version-specific limitations.
- Implement data quality and validation components to ensure data security, integrity, and consistency across pipelines.
- Identify and address schema drift, data changes, and backfill requirements within data pipelines.
- Troubleshoot pipeline failures, performance issues, and data inconsistencies across distributed AWS environments.
- Collaborate with data engineers, architects, and other technical stakeholders to develop reliable and maintainable data solutions.
- Contribute to technical design decisions, documentation, testing, and continuous improvement of the data platform.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Qualifications:
\n
- \n
- Strong hands-on experience with Java and/or Scala.
- Strong experience with Apache Spark and distributed data processing.
- Hands-on experience building AWS data pipelines using:
- Amazon S3
- AWS Glue
- AWS Step Functions
- AWS Lambda
- Experience with event-driven architectures, particularly EventBridge rules and S3 event notifications.
- Understanding of Step Functions state machines, scheduling, orchestration, and downstream job synchronization.
- Experience configuring and scaling AWS Glue jobs.
- Strong understanding of data quality, schema validation, schema drift, and backfill strategies.
- Experience troubleshooting and supporting production data pipelines.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Preferred Qualifications
\n
- \n
- Experience with Apache Iceberg, including Glue/Iceberg integration and troubleshooting.
- Experience designing resilient pipelines with rerun, retry, and failure-recovery mechanisms.
- Experience implementing data validation or security-focused data quality components.
- Experience working with large-scale cloud data platforms and distributed processing environments.
- Strong problem-solving skills and ability to independently research and resolve technical limitations.
\n
\n
\n
\n
\n
\n
