Search

Site Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Chicago, IL (Hybrid – 3X a week)

\n


\n

About NOCD

\n

NOCD is the #1 telehealth provider for the treatment of obsessive-compulsive disorder (OCD). Through our technology platform, we connect Members with licensed Therapists who specialize in Exposure and Response Prevention (ERP), the gold-standard treatment for OCD.

\n


\n

We're building technology that helps people reclaim their lives and we're growing quickly. Our platform supports Members, Therapists, and teams across a nationwide healthcare organization, creating unique engineering challenges around scale, reliability, security, real-time communication, and healthcare infrastructure.

\n


\n

The Opportunity

\n

Build the infrastructure behind a rapidly growing healthcare platform.

\n

NOCD is looking for a Senior Site Reliability Engineer (SRE) to help shape the infrastructure, reliability, and developer platform powering our next stage of growth.

\n


\n

This is not a traditional DevOps or infrastructure-maintenance role. You'll have meaningful ownership over how we build, deploy, monitor, secure, and scale our systems.

\n


\n

You'll work across AWS, Kubernetes, Terraform, CI/CD, observability, security, and software engineering, partnering closely with application engineers to make our systems more reliable while helping our teams ship faster.

\n


\n

What You'll Own

\n

    \n
  • Design, build, and evolve AWS infrastructure supporting NOCD's production applications.
  • \n

  • Build scalable infrastructure using Terraform, Kubernetes, and Docker.
  • \n

  • Improve CI/CD, deployment automation, and developer tooling to help engineers ship faster and more safely.
  • \n

  • Own the reliability, availability, performance, and scalability of production systems.
  • \n

  • Build observability across infrastructure and applications, including monitoring, logging, tracing, and alerting.
  • \n

  • Establish SLIs, SLOs, and operational metrics and lead incident response and root-cause analysis.
  • \n

  • Write production-quality software in Python, TypeScript, or similar languages.
  • \n

  • Build automation, APIs, microservices, and internal engineering tools.
  • \n

  • Partner with Engineering and Security on HIPAA, SOC 2, infrastructure security, and compliance.
  • \n

  • Own technical initiatives from architecture through production and help shape engineering standards and best practices.
  • \n

\n


\n

What We're Looking For

\n

    \n
  • 7+ years of professional software engineering experience, with significant experience in SRE, platform engineering, DevOps, or cloud infrastructure.
  • \n

  • Strong hands-on experience with AWS and cloud architecture.
  • \n

  • Strong experience with Terraform or other infrastructure-as-code tools.
  • \n

  • Strong experience with Docker and Kubernetes.
  • \n

  • Proficiency in Python, TypeScript, or a similar programming language.
  • \n

  • Experience designing and operating CI/CD pipelines and deployment automation.
  • \n

  • Strong understanding of distributed systems, APIs, networking, databases, and cloud architecture.
  • \n

  • Experience with observability, production troubleshooting, and incident response.
  • \n

  • Strong understanding of reliability, scalability, availability, and performance.
  • \n

  • Ability to own technical initiatives and work effectively across engineering teams.
  • \n

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field, or equivalent professional experience.
  • \n

\n


\n

Nice to Have

\n

    \n
  • Experience in healthcare, fintech, or another regulated environment.
  • \n

  • Experience with HIPAA, SOC 2, HITRUST, or similar compliance frameworks.
  • \n

  • Experience with Datadog, CloudWatch, Prometheus, Grafana, Splunk, or similar tools.
  • \n

  • Experience with GitHub Actions, Jenkins, ArgoCD, or GitOps.
  • \n

  • Experience designing highly available or multi-region AWS architectures.
  • \n

  • Experience with DevSecOps, IAM, secrets management, and encryption.
  • \n

  • Experience building internal developer platforms or tooling.
  • \n

  • Experience with disaster recovery, capacity planning, or performance engineering.
  • \n

\n


\n

What Success Looks Like

\n

    \n
  • Build more reliable, observable, and scalable infrastructure.
  • \n

  • Make it easier for engineers to ship quickly and safely.
  • \n

  • Strengthen our cloud, deployment, and security architecture.
  • \n

  • Establish stronger reliability and incident-management practices.
  • \n

  • Build the technical foundation needed for NOCD's continued growth.
  • \n

\n


\n

What We Offer

\n

    \n
  • Comprehensive benefits package, including medical, dental, vision coverage, and 401(k) match
  • \n

  • 11 observed company holidays per year
  • \n

  • PTO based on an accrual system
  • \n

  • Downtown Chicago office with an on-site gym
  • \n

  • Engaging, mission-driven startup environment
  • \n

  • NOCD provides 12 weeks of fully paid parental leave for the primary caregiver and 6 weeks for the secondary caregiver, for qualifying full-time employees.
  • \n

\n


\n

Pay Transparency

\n

The expected pay range for this position is $150,000 to $200,000. Actual pay will be based on the individual's qualifications and experience. This role is also eligible for annual performance-based incentives tied to individual achievement and company-wide goals.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...