Search

Senior Site Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

🔹 Position Overview

\n

We are looking for a highly experienced Senior Site Reliability Engineer (SRE) with strong hands-on experience in Oracle Transportation Management (OTM).

\n

The ideal candidate will have 10+ years of overall IT/SRE experience, including at least 5 years of dedicated hands-on OTM experience, with strong expertise in OTM production environments, integrations, deployments, troubleshooting, monitoring, automation, and reliability engineering.

\n

The role will focus on designing, implementing, and maintaining highly available, scalable, secure, and reliable OTM environments across on-premises and OTM SaaS/cloud platforms using SRE best practices.

\n


\n

🔹 Key Responsibilities

\n

    \n
  • Design, implement, maintain, and improve highly available and scalable Oracle Transportation Management (OTM) environments.
  • \n

  • Support both on-premises OTM and OTM SaaS/cloud environments.
  • \n

  • Perform SRE activities including monitoring, incident management, problem management, capacity planning, performance optimization, automation, and reliability engineering.
  • \n

  • Monitor OTM application health, infrastructure performance, integrations, and critical business processes using observability and monitoring tools.
  • \n

  • Troubleshoot complex issues across OTM applications, middleware, databases, integrations, infrastructure, and cloud environments.
  • \n

  • Drive production incidents through resolution while minimizing business impact.
  • \n

  • Perform Root Cause Analysis (RCA) and implement long-term reliability improvements.
  • \n

  • Support OTM SaaS releases, patching, configuration, deployments, upgrades, migrations, and environment management.
  • \n

  • Work closely with application, development, infrastructure, database, integration, and cloud teams.
  • \n

  • Leverage Snowflake Observe and other observability capabilities to identify trends, detect anomalies, and improve system performance.
  • \n

  • Define and maintain SLOs, SLIs, alerting, dashboards, and reliability metrics for critical OTM services.
  • \n

  • Apply SRE best practices across automation, observability, security, compliance, capacity management, and disaster recovery.
  • \n

  • Support cloud-based OTM SaaS solutions and contribute to cloud engineering initiatives.
  • \n

  • Identify opportunities for automation and eliminate repetitive operational activities.
  • \n

\n


\n

🔹 Required Qualifications

\n

    \n
  • 10+ years of overall IT/SRE experience.
  • \n

  • Minimum 5 years of dedicated hands-on Oracle Transportation Management (OTM) experience.
  • \n

  • Strong hands-on knowledge of OTM application, production support, integrations, deployments, and troubleshooting.
  • \n

  • Experience working with OTM in both on-premises and SaaS/cloud environments.
  • \n

  • Strong knowledge of Oracle Database.
  • \n

  • Strong knowledge of Java programming.
  • \n

  • Hands-on experience with:
  • \n

  • Incident Response
  • \n

  • Monitoring & Observability
  • \n

  • Root Cause Analysis
  • \n

  • Performance Tuning
  • \n

  • Automation
  • \n

  • Capacity Planning
  • \n

  • Reliability Engineering
  • \n

  • Strong scripting/automation experience with Python, Bash, Shell, or similar technologies.
  • \n

  • Experience with OTM SaaS migrations, upgrades, releases, patching, and environment transitions.
  • \n

  • Experience with cloud/SaaS platforms, preferably Oracle Cloud Infrastructure (OCI) and Oracle OTM SaaS.
  • \n

  • Experience with observability platforms such as Snowflake Observe, Oracle Enterprise Manager (OEM), or equivalent.
  • \n

  • Experience with Infrastructure as Code (IaC) or configuration management tools such as Terraform, Puppet, or Ansible.
  • \n

\n


\n

🔹 Preferred / Additional Skills

\n

    \n
  • Google Cloud Platform (GCP) experience.
  • \n

  • Cloud-native technologies.
  • \n

  • Docker / Kubernetes.
  • \n

  • Java and .NET applications.
  • \n

  • Kafka.
  • \n

  • REST/SOAP APIs.
  • \n

  • Enterprise integration technologies.
  • \n

  • Relational databases and distributed systems.
  • \n

  • Security, compliance, and disaster recovery.
  • \n

  • SRE frameworks and reliability metrics.
  • \n

\n


\n

🔹 Education

\n

Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent professional experience.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...