Search

321,671 Jobs

No logo available
Taylor Farms CO
location2713 W Cucharras St, Colorado Springs, CO 80904, USA
PublishedPublished: 7/24/2026
No logo available
Turnkey Corrections
locationRiver Falls, WI 54022, USA
PublishedPublished: 7/24/2026
No logo available
Unique Cleaning Service, Inc.
locationAppleton, WI, USA
PublishedPublished: 7/24/2026
No logo available
Aptive Environmental LLC
locationMurfreesboro, TN, USA
PublishedPublished: 7/24/2026
No logo available
Butterfly Network
locationBurlington, MA, USA
PublishedPublished: 7/24/2026
No logo available
Emerge Recovery & Trade
locationXenia, OH 45385, USA
PublishedPublished: 7/24/2026
No logo available
Tarana Wireless, Inc.
locationMilpitas, CA 95035, USA
PublishedPublished: 7/24/2026
No logo available
E-Space
locationSaratoga, CA, USA
PublishedPublished: 7/24/2026
No logo available
Atrium
locationNew York, NY, USA
PublishedPublished: 7/24/2026
No logo available
hth companies
locationUnion, MO 63084, USA
PublishedPublished: 7/24/2026

Site Reliability Engineer

Coforge
locationSeattle, WA, USA
PublishedPublished: 6/14/2022
Technology
Full Time

Job Description

Job Title: Site Reliability Engineer (SRE) - Middleware API

\n

Key Skills: SRE, Middleware API, Scripting and Automation tools

\n

Experience: 10+ Years’ experience

\n

Location: Seattle, Washington & Atlanta, Georgia

\n


\n

We at Coforge are hiring experienced professionals with strong knowledge of Site Reliability Engineering (SRE), middleware APIs, scripting (Python/Bash/Ansible), cloud platforms (AWS/Azure/GCP), Kubernetes, monitoring tools (Prometheus, Grafana, ELK), and incident management.

\n


\n

Key Responsibilities:

\n

    \n
  • Drive site reliability engineering (SRE) practices to ensure high availability, scalability, and performance of systems.
  • \n

  • Design, develop, and maintain middleware APIs and integrations.
  • \n

  • Build and manage CI/CD pipelines with automation using tools and scripting (Python, Bash, Ansible).
  • \n

  • Deploy, manage, and optimize applications on cloud platforms (AWS/Azure/GCP).
  • \n

  • Work with Kubernetes for container orchestration, deployment, and scaling.
  • \n

  • Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK).
  • \n

  • Lead incident management, including troubleshooting, root cause analysis (RCA), and resolution.
  • \n

  • Ensure system reliability, uptime, and performance through proactive measures and automation.
  • \n

  • Collaborate with development teams to improve system design, resiliency, and observability.
  • \n

  • Enforce best practices for security, compliance, and operational excellence.
  • \n

\n


\n

Required Skills

\n

    \n
  • Design, implement, and maintain robust Middleware API solutions to ensure high availability and performance.
  • \n

  • Monitor and troubleshoot system performance, identifying and resolving issues proactively.
  • \n

  • Collaborate with development teams to integrate SRE practices into the software development lifecycle.
  • \n

  • Develop and maintain automation tools for deployment, monitoring, and incident response.
  • \n

  • Establish and enforce best practices for API design, security, and documentation.
  • \n

  • Participate in on-call rotations and incident response efforts to ensure system reliability.
  • \n

  • Conduct post-incident reviews and implement improvements based on findings.
  • \n

  • Provide technical guidance and mentorship to junior team members.
  • \n

\n


\n

Preferred Qualifications

\n

    \n
  • Strong SRE knowledge, middleware API experience, scripting (Python/Bash/Ansible), cloud (AWS/Azure/GCP), Kubernetes, monitoring (Prometheus/Grafana/ELK), and incident management expertise.
  • \n

\n