Job Description
Our client is a growing technology-driven organization in the financial services space that is modernizing a traditionally legacy-heavy industry through cloud technology, automation, data, and AI.
\n
\n
They are looking for a Lead Site Reliability Engineer (SRE) to take a hands-on leadership role in building and maintaining reliable, scalable, and secure cloud infrastructure. This person will lead a team of SREs while also remaining technically involved in infrastructure, automation, observability, incident response, and continuous improvement.
\n
This is an opportunity to have a meaningful impact on the reliability and scalability of a modern AWS environment while helping establish strong engineering and operational practices.
\n
\n
What You'll Do
\n
- \n
- Lead a team of SREs, including full-time employees and contractors
- Assign technical work, review pull requests, and provide technical guidance and feedback
- Design, implement, and maintain scalable, reliable, and secure infrastructure in AWS
- Build and maintain monitoring, alerting, logging, and dashboarding solutions
- Automate infrastructure provisioning and configuration using Terraform and Terragrunt
- Develop and maintain CI/CD pipelines to improve deployment reliability and developer productivity
- Support and enhance Kubernetes-based environments and containerized workloads
- Implement and manage Argo CD and Argo Workflows for application and workflow orchestration
- Respond to production incidents, lead root cause analysis, and implement long-term solutions
- Identify opportunities to improve system performance, reliability, security, and cost efficiency
- Partner with software engineering and other technical teams to improve infrastructure and developer workflows
- Establish and maintain disaster recovery strategies and processes
- Promote infrastructure, security, automation, and operational best practices across engineering
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
What We're Looking For
\n
\n
Required Qualifications
\n
- \n
- Bachelor's degree in Computer Science or a related field, or equivalent professional experience
- Strong hands-on experience with AWS, including services such as EKS, Fargate, and Aurora
- Strong experience with Kubernetes and related tooling/plugins
- Hands-on experience with Terraform and Terragrunt
- Experience with Argo CD and Argo Workflows
- Experience building and maintaining CI/CD pipelines, particularly with GitHub Actions
- Strong understanding of Docker and containerization
- Experience with cloud infrastructure security and compliance
- Experience with logging, monitoring, and observability platforms
- Strong scripting/programming experience
- Experience with database management
- Proficiency with Git/version control
- Experience with production incident management and troubleshooting
- Previous experience leading or mentoring technical engineers
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Preferred Qualifications
\n
- \n
- Experience with Datadog
- Experience with Cloudflare
- Experience implementing advanced AWS security capabilities such as GuardDuty and Security Hub
- Experience developing and maintaining disaster recovery strategies
- Experience with workflow automation platforms such as Camunda
- Experience within mortgage, lending, banking, or financial services
- AWS certifications such as AWS Certified Developer/DevOps Engineer or Associate-level certifications
- Master's degree in Computer Science or a related field
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
What You'll Get
\n
- \n
- Competitive base salary and performance-based compensation
- Comprehensive health, dental, and vision benefits
- 401(k) with company matching
- Opportunities for career growth as the organization scales
- Collaborative, technology-focused engineering environment
- Generous PTO and additional employee perks
\n
\n
\n
\n
\n
\n
\n
\n
This is a hands-on technical leadership opportunity for an SRE who enjoys building reliable cloud infrastructure, improving engineering practices, and leading a team while remaining close to the technology.
