Search

307,253 Jobs

No logo available
innovitusa
locationJackson, MS, USA
PublishedPublished: 7/18/2026
No logo available
Express Employment Professionals
locationHutchinson, KS, USA
PublishedPublished: 7/18/2026
No logo available
Xylem I LLC
locationMonroe, NC, USA
PublishedPublished: 7/18/2026
No logo available
Sierra Gold Nurseries
locationYuba City, CA, USA
PublishedPublished: 7/18/2026
No logo available
Allied Universal
locationAlexandria, VA, USA
PublishedPublished: 7/18/2026
No logo available
Oakley Services
locationFairburn, GA, USA
PublishedPublished: 7/18/2026
No logo available
Softpath
locationSunrise, FL, USA
PublishedPublished: 7/18/2026
No logo available
innovitusa
locationDenver, CO, USA
PublishedPublished: 7/18/2026
No logo available
Strategic Resources International USA
locationFarmington, MI, USA
PublishedPublished: 7/18/2026
No logo available
Wonder Group, INC
locationBoston, MA, USA
PublishedPublished: 7/18/2026

Site Reliability Engineer III

Robert Half
locationMt Laurel Township, NJ, USA
PublishedPublished: 6/14/2022
Technology
Full Time

Job Description

Job Description

Site Reliability Engineer III

Location: Onsite in Mount Laurel, NJ

Duration: Through 12/31/2026, extensions likely


We are seeking a Site Reliability Engineer (SRE) to support cloud infrastructure, automation, reliability, security, and observability initiatives for AI/ML platform environments. This role is ideal for a mid-to-senior-level engineer with strong cloud and platform engineering experience who can drive scalability, reliability, and automation within Kubernetes-based production environments.


Responsibilities:

  • Support Site Reliability Engineering initiatives across AI/ML platform environments.
  • Deploy, maintain, and optimize cloud infrastructure across AWS and GCP.
  • Build, manage, and maintain Infrastructure as Code (IaC) solutions using Terraform.
  • Improve platform reliability, scalability, security, and operational efficiency.
  • Administer and support Kubernetes and Amazon EKS environments.
  • Monitor and troubleshoot system performance using observability and monitoring tools including Prometheus, Grafana, Datadog, and Elasticsearch.
  • Automate operational processes, workflows, and routine administrative tasks using Python and related tooling.
  • Support and enhance CI/CD pipelines and deployment automation.
  • Troubleshoot complex distributed systems and production issues in highly available environments.
  • Collaborate with engineering teams to improve platform performance, monitoring, and operational resiliency.
  • Work with technologies including Kubernetes, Docker, AWS, GCP, EKS, Terraform, Prometheus, Grafana, Datadog, Elasticsearch, MySQL, Kafka, and Python.

Qualifications:

  • 4–8 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or related disciplines.
  • Strong hands-on experience with AWS cloud services.
  • Hands-on Kubernetes administration and support experience.
  • Expertise with Terraform and Infrastructure as Code (IaC) practices.
  • Experience designing, supporting, and improving CI/CD pipelines.
  • Strong observability and monitoring experience, particularly with Prometheus.
  • Experience with Grafana, Datadog, Elasticsearch, or similar monitoring platforms.
  • Python programming and scripting proficiency.
  • Experience supporting distributed systems and large-scale, highly available production environments.
  • Knowledge of algorithms, data structures, software design principles, and system troubleshooting.
  • Experience with Docker containers and cloud-native platforms.

Preferred Qualifications:

  • Bachelor's degree in Computer Science or a related technical discipline.
  • Experience supporting AI/ML platforms or infrastructure.