Search

Lead Site Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Our client is a growing technology-driven organization in the financial services space that is modernizing a traditionally legacy-heavy industry through cloud technology, automation, data, and AI.

\n


\n

They are looking for a Lead Site Reliability Engineer (SRE) to take a hands-on leadership role in building and maintaining reliable, scalable, and secure cloud infrastructure. This person will lead a team of SREs while also remaining technically involved in infrastructure, automation, observability, incident response, and continuous improvement.

\n

This is an opportunity to have a meaningful impact on the reliability and scalability of a modern AWS environment while helping establish strong engineering and operational practices.

\n


\n

What You'll Do

\n

    \n
  • Lead a team of SREs, including full-time employees and contractors
  • \n

  • Assign technical work, review pull requests, and provide technical guidance and feedback
  • \n

  • Design, implement, and maintain scalable, reliable, and secure infrastructure in AWS
  • \n

  • Build and maintain monitoring, alerting, logging, and dashboarding solutions
  • \n

  • Automate infrastructure provisioning and configuration using Terraform and Terragrunt
  • \n

  • Develop and maintain CI/CD pipelines to improve deployment reliability and developer productivity
  • \n

  • Support and enhance Kubernetes-based environments and containerized workloads
  • \n

  • Implement and manage Argo CD and Argo Workflows for application and workflow orchestration
  • \n

  • Respond to production incidents, lead root cause analysis, and implement long-term solutions
  • \n

  • Identify opportunities to improve system performance, reliability, security, and cost efficiency
  • \n

  • Partner with software engineering and other technical teams to improve infrastructure and developer workflows
  • \n

  • Establish and maintain disaster recovery strategies and processes
  • \n

  • Promote infrastructure, security, automation, and operational best practices across engineering
  • \n

\n


\n

What We're Looking For

\n


\n

Required Qualifications

\n

    \n
  • Bachelor's degree in Computer Science or a related field, or equivalent professional experience
  • \n

  • Strong hands-on experience with AWS, including services such as EKS, Fargate, and Aurora
  • \n

  • Strong experience with Kubernetes and related tooling/plugins
  • \n

  • Hands-on experience with Terraform and Terragrunt
  • \n

  • Experience with Argo CD and Argo Workflows
  • \n

  • Experience building and maintaining CI/CD pipelines, particularly with GitHub Actions
  • \n

  • Strong understanding of Docker and containerization
  • \n

  • Experience with cloud infrastructure security and compliance
  • \n

  • Experience with logging, monitoring, and observability platforms
  • \n

  • Strong scripting/programming experience
  • \n

  • Experience with database management
  • \n

  • Proficiency with Git/version control
  • \n

  • Experience with production incident management and troubleshooting
  • \n

  • Previous experience leading or mentoring technical engineers
  • \n

\n


\n

Preferred Qualifications

\n

    \n
  • Experience with Datadog
  • \n

  • Experience with Cloudflare
  • \n

  • Experience implementing advanced AWS security capabilities such as GuardDuty and Security Hub
  • \n

  • Experience developing and maintaining disaster recovery strategies
  • \n

  • Experience with workflow automation platforms such as Camunda
  • \n

  • Experience within mortgage, lending, banking, or financial services
  • \n

  • AWS certifications such as AWS Certified Developer/DevOps Engineer or Associate-level certifications
  • \n

  • Master's degree in Computer Science or a related field
  • \n

\n


\n

What You'll Get

\n

    \n
  • Competitive base salary and performance-based compensation
  • \n

  • Comprehensive health, dental, and vision benefits
  • \n

  • 401(k) with company matching
  • \n

  • Opportunities for career growth as the organization scales
  • \n

  • Collaborative, technology-focused engineering environment
  • \n

  • Generous PTO and additional employee perks
  • \n

\n


\n

This is a hands-on technical leadership opportunity for an SRE who enjoys building reliable cloud infrastructure, improving engineering practices, and leading a team while remaining close to the technology.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...