Search

Platform Director

PublishedPublished: 6/14/2022
Real Estate

Job Description

Job Title: Director of Platform Engineering

\n

Location: Chicago, IL (Onsite – Hybrid - 3 days in a week)

\n

Position Type: Fulltime permanent position

\n


\n

Responsibilities:

\n

Notes – Even though the primary role of this person will be to manage a team of engineers and managers, they need to have a strong technical background in SRE, Observability, Infrastructure, Kubernetes, Kafka, etc.

\n


\n

2-3 sentences describing the overall responsibilities.

\n

This is a people-management role with full supervisory responsibility for a team of engineers and managers. The Director will lead and manage this team to drive reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing environment. The ideal candidate is a hands-on, data-driven technical leader, who can personally raise the bar on SRE and observability practices, reduce waste, optimize cloud efficiency, and improve performance, in close partnership with a dedicated SRE/monitoring team and a centralized architecture function. This individual will also serve as a player-coach, passionate about new technologies, mentoring technical teams through complex initiatives while keeping a steady eye on the extensive regulatory/compliance demands on our company (e.g., CIS, NIST).

\n


\n

Qualifications:

\n

The requirements listed are representative of the knowledge, skill, and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the primary functions.

\n

    \n
  • [Required] 5+ years of demonstrated experience leading engineering teams, with an emphasis on developing key talent and cultivating positive, high-performing cultures
  • \n

  • [Required] 10+ years of progressive, hands-on experience in software engineering with an understanding of large-scale computing solutions (primarily AWS), including software design and development, database architectures, IP networking, security, cloud operations, and performance tuning
  • \n

  • [Required] Demonstrated, hands-on expertise building and maturing SRE and observability practices at scale, able to personally raise the technical bar for a dedicated SRE/monitoring team, not just consume their output
  • \n

  • [Required] Demonstrated track record owning the definition and governance of SLOs and SLAs for mission-critical systems, and personally driving resilience engineering practices (chaos engineering, load/performance testing) to validate reliability ahead of incidents
  • \n

  • [Required] Demonstrated track record driving reliability through a data-driven culture, waste/toil reduction, cloud efficiency measures, and performance optimization, treating metrics as the primary lens for every decision
  • \n

  • [Required] Experience defining, instrumenting, and acting on software delivery performance metrics (DORA metrics, cycle time, deployment frequency, lead time, etc.) to drive engineering improvement initiatives
  • \n

  • [Required] Strong consultative, communication, team player, and analytical skills, with the ability to regularly interact between various teams distributed across the US
  • \n

  • [Required] Strong technical team leadership and technical project management skills
  • \n

  • [Required] Relevant experience leading highly technical team members through adopting new technologies while maintaining highly available, mission-critical systems, with a proven track record of success
  • \n

  • [Required] Ability to clearly communicate verbally and in writing to business and technology leaders, architects, developers, and team members
  • \n

  • [Required] Must be able to collaborate effectively with a group of high-performing, technical individuals
  • \n

  • [Required] Experience acting as a product owner, defining roadmap, requirements, and priorities, for a platform capability such as observability, ideally in partnership with a separate team that owns the underlying tooling and operations
  • \n

  • [Required] Experience with architecting, implementing, and maintaining highly available mission-critical environments for 24x7 availability
  • \n

  • [Required] Demonstrated history of working within deadlines and ability to work well under pressure
  • \n

  • [Required] Experience managing work tasks using Agile methodology/scrum desired
  • \n

  • [Preferred] Comfort with ambiguity and demonstrated ability to lead complex programs in a decentralized environment
  • \n

  • [Preferred] Experience working in an environment with a defined production change control process; experience working with audits and compliance or in a regulated environment a plus
  • \n

  • [Preferred] Experience in organizations with a mature, centralized SRE function, avoiding role/scope overlap
  • \n

\n


\n

Technical Skills:

\n

    \n
  • [Required] Deep expertise in OpenTelemetry, including instrumentation standards, auto-instrumentation, semantic conventions, and the OTel Collector, as the foundation for a vendor-neutral, paved-road instrumentation strategy
  • \n

  • [Required] Hands-on experience with metrics engines and time-series databases, including Prometheus and PromQL, plus scale-out options such as Mimir, Thanos, VictoriaMetrics, or Amazon Managed Prometheus, including cardinality management
  • \n

  • [Required] Hands-on experience with tracing and logging backends such as Tempo/Jaeger, Loki/Elastic/Splunk (Splunk is common in financial services), and AWS X-Ray
  • \n

  • [Required] Experience with Kubernetes and Kafka observability specifically, including EKS metrics, kube-state-metrics, consumer lag, and broker health
  • \n

  • [Required] Experience with alerting and incident tooling such as PagerDuty or Opsgenie, including ServiceNow integration, alert routing, and noise reduction
  • \n

  • [Required] Hands-on experience with SLO-as-code frameworks such as OpenSLO, Sloth, or Nobl9, defining and governing SLOs in the repo rather than only in a vendor UI
  • \n

  • [Required] Experience delivering golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box
  • \n

  • [Required] Experience with resilience validation practices, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing
  • \n

  • [Required] Deep, hands-on mastery of observability tooling and practices (metrics, distributed tracing, centralized logging, dashboards, alerting), e.g., Datadog, Prometheus/Grafana, CloudWatch, or equivalent, sufficient to elevate, not just consume, a dedicated SRE team’s capability
  • \n

  • [Required] Deep understanding of SRE principles including SLOs/SLIs, error budgets, incident management, and postmortem culture, with a track record of driving adoption and maturity
  • \n

  • [Required] Fluency in the Golden Signals, RED, and USE methods as applied frameworks for monitoring and alerting design
  • \n

  • [Required] Hands-on experience with: Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, Github, and configuration management tools such as Puppet, Chef, or Ansible
  • \n

  • [Required] Deep, hands-on expertise with infrastructure-as-code (IaC) tools and practices (e.g., Terraform, CloudFormation, CDK, Pulumi), with a track record of driving IaC adoption at scale across a cloud platform organization
  • \n

  • [Required] Relevant experience with configuration and implementation of IaaS, Infrastructure as Code, AWS, Azure, etc.
  • \n

  • [Required] Expert working knowledge of infrastructure design and components, such as servers, operating systems, networks, and storage
  • \n

  • [Required] Basic understanding of good delivery practices and continual integration and improvement; Agile/Lean background for projects and project delivery
  • \n

  • [Preferred] Experience with telemetry pipeline tools such as Fluent Bit, Vector, or Cribl for routing, sampling, redaction, and cost control
  • \n

  • [Preferred] Experience with synthetic monitoring, real user monitoring (RUM), and eBPF-based observability
  • \n

  • [Preferred] Familiarity with audit-grade telemetry practices (log retention/immutability, PII and sensitive-data scrubbing, access controls on observability data) and mapping observability controls to compliance frameworks such as CIS, NIST, or SIFMU-specific resilience requirements
  • \n

  • [Preferred] Familiarity with engineering metrics/analytics platforms (e.g., LinearB, Jellyfish, Sleuth, Haystack, or internal equivalents) used to track DORA metrics and delivery performance
  • \n

  • [Preferred] Competent in all phases of application development and implementation, including SDLC; hands-on scripting/development skills in Python, Ruby, Go, Java, etc. in a corporate environment strongly desired
  • \n

  • [Preferred] Experience establishing IaC governance and standards (module libraries, policy-as-code, drift detection) across multiple teams or business units
  • \n

  • [Preferred] Experience building a metrics-driven engineering culture, including scorecards, dashboards, or leadership reporting on delivery performance
  • \n

\n


\n

Education and/or Experience:

\n

    \n
  • [Required] Bachelor’s degree, preferably in a technical discipline (Computer Science, Mathematics, etc.), or equivalent combination of education and experience required; Master’s degree and relevant experience also considered
  • \n

  • [Required] 10+ years’ experience in IT systems installation, operations, administration, and maintenance of cloud systems / virtualized servers, including 5+ years in a technical leadership role
  • \n

  • [Preferred] Experience working in a financial services or highly regulated environment preferred
  • \n

\n


\n

Certificates or Licenses:

\n

    \n
  • [Required] AWS Solutions Architect Associate Certification or higher strongly desired
  • \n

  • [Preferred] Relevant industry certifications such as Microsoft Azure or Google Cloud
  • \n

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...