Job Description
Job DescriptionReliability Engineer
Location: Reston, Virginia
Work Arrangement: Remote/Hybrid – Onsite in Reston, VA may be requested
Duration: 6 Months
Position Overview
We are seeking an experienced Reliability Engineer with strong expertise in AWS cloud environments, multi-region workloads, resiliency, and failover optimization.
The ideal candidate will have hands-on experience designing, supporting, and optimizing highly available AWS infrastructure across multiple regions, with a strong understanding of Infrastructure as Code and automation.
Key Responsibilities
- Design, implement, and optimize reliability and resiliency strategies for multi-region AWS workloads.
- Analyze and improve AWS infrastructure failover mechanisms and recovery processes.
- Support high availability, disaster recovery, and business continuity initiatives.
- Develop and maintain Infrastructure as Code using AWS CloudFormation and Terraform.
- Build, maintain, and enhance AWS automation for resiliency and failover.
- Work with serverless AWS services including Lambda and Step Functions.
- Support and troubleshoot AWS messaging and storage services, including AWS MQ and AWS EFS.
- Identify potential reliability risks, bottlenecks, and single points of failure.
- Test and validate failover and recovery procedures.
- Collaborate with engineering and cloud teams to improve system availability and operational resilience.
- Modify existing automation code to address reliability and resiliency requirements.
Required AWS Skills
- Strong experience with AWS cloud environments.
- Experience supporting multi-region AWS workloads.
- Strong understanding of AWS infrastructure resiliency and failover.
- Hands-on experience with Infrastructure as Code (IaC).
- Strong experience with:
- AWS CloudFormation
- CloudFormation YAML
- Terraform
- AWS Lambda
- AWS Step Functions
- AWS MQ
- AWS EFS
Programming Requirements
- Strong Python experience.
- Ability to understand existing Python-based resiliency and automation code.
- Ability to modify, troubleshoot, and enhance existing automation scripts/code.
- Understanding of automation and scripting practices for cloud infrastructure.
