Director Platform Engineering - SRE / Observability
$175k - $240kRequest Technology, LLC
Director, Platform Engineering
SALARY: $175k - $240k plus 30% bonus
LOCATION: CHICAGO, IL
HYBRID 3 DAYS ONSITE
You will manage a team of SREs and managers. 7-10 people and managers. Keys SRE observability infrastructure kubernetes kafka aws terraform. j
The Director will lead and manage this team to drive reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing environment. The ideal candidate is a hands-on, data-driven technical leader, who can personally raise the bar on SRE and observability practices, reduce waste, optimize cloud efficiency, and improve performance, in close partnership with a dedicated SRE/monitoring team and a centralized architecture function.
Qualifications:
- [Required] 5+ years of demonstrated experience leading engineering teams, with an emphasis on developing key talent and cultivating positive, high-performing cultures
- [Required] 10+ years of progressive, hands-on experience in software engineering with an understanding of large-scale computing solutions (primarily AWS), including software design and development, database architectures, IP networking, security, cloud operations, and performance tuning
- [Required] Demonstrated, hands-on expertise building and maturing SRE and observability practices at scale, able to personally raise the technical bar for a dedicated SRE/monitoring team, not just consume their output
- [Required] Demonstrated track record owning the definition and governance of SLOs and SLAs for mission-critical systems, and personally driving resilience engineering practices (chaos engineering, load/performance testing) to validate reliability ahead of incidents
- [Required] Demonstrated track record driving reliability through a data-driven culture, waste/toil reduction, cloud efficiency measures, and performance optimization, treating metrics as the primary lens for every decision
- [Required] Experience defining, instrumenting, and acting on software delivery performance metrics (DORA metrics, cycle time, deployment frequency, lead time, etc.) to drive engineering improvement initiatives
- [Required] Strong consultative, communication, team player, and analytical skills, with the ability to regularly interact between various teams distributed across the US
- [Required] Strong technical team leadership and technical project management skills
- [Required] Relevant experience leading highly technical team members through adopting new technologies while maintaining highly available, mission-critical systems, with a proven track record of success
- [Required] Ability to clearly communicate verbally and in writing to business and technology leaders, architects, developers, and team members
- [Required] Must be able to collaborate effectively with a group of high-performing, technical individuals
- [Required] Experience acting as a product owner, defining roadmap, requirements, and priorities, for a platform capability such as observability, ideally in partnership with a separate team that owns the underlying tooling and operations
- [Required] Experience with architecting, implementing, and maintaining highly available mission-critical environments for 24x7 availability
- [Required] Demonstrated history of working within deadlines and ability to work well under pressure
- [Required] Experience managing work tasks using Agile methodology/scrum desired
Technical Skills:
- [Required] Deep expertise in OpenTelemetry, including instrumentation standards, auto-instrumentation, semantic conventions, and the OTel Collector, as the foundation for a vendor-neutral, paved-road instrumentation strategy
- [Required] Hands-on experience with metrics engines and time-series databases, including Prometheus and PromQL, plus scale-out options such as Mimir, Thanos, VictoriaMetrics, or Amazon Managed Prometheus, including cardinality management
- [Required] Hands-on experience with tracing and logging backends such as Tempo/Jaeger, Loki/Elastic/Splunk (Splunk is common in financial services), and AWS X-Ray
- [Required] Experience with Kubernetes and Kafka observability specifically, including EKS metrics, kube-state-metrics, consumer lag, and broker health
- [Required] Experience with alerting and incident tooling such as PagerDuty or Opsgenie, including ServiceNow integration, alert routing, and noise reduction
- [Required] Hands-on experience with SLO-as-code frameworks such as OpenSLO, Sloth, or Nobl9, defining and governing SLOs in the repo rather than only in a vendor UI
- [Required] Experience delivering golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box
- [Required] Experience with resilience validation practices, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing
- [Required] Deep, hands-on mastery of observability tooling and practices (metrics, distributed tracing, centralized logging, dashboards, alerting), e.g., Datadog, Prometheus/Grafana, CloudWatch, or equivalent, sufficient to elevate, not just consume, a dedicated SRE team’s capability
- [Required] Deep understanding of SRE principles including SLOs/SLIs, error budgets, incident management, and postmortem culture, with a track record of driving adoption and maturity
- [Required] Fluency in the Golden Signals, RED, and USE methods as applied frameworks for monitoring and alerting design
- [Required] Hands-on experience with: Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, Github, and configuration management tools such as Puppet, Chef, or Ansible
- [Required] Deep, hands-on expertise with infrastructure-as-code (IaC) tools and practices (e.g., Terraform, CloudFormation, CDK, Pulumi), with a track record of driving IaC adoption at scale across a cloud platform organization
- [Required] Relevant experience with configuration and implementation of IaaS, Infrastructure as Code, AWS, Azure, etc.
- [Required] Expert working knowledge of infrastructure design and components, such as servers, operating systems, networks, and storage
- [Required] Basic understanding of good delivery practices and continual integration and improvement; Agile/Lean background for projects and project delivery
- [Required] Bachelor’s degree, preferably in a technical discipline (Computer Science, Mathematics, etc.), or equivalent combination of education and experience required; Master’s degree and relevant experience also considered
- [Required] 10+ years’ experience in IT systems installation, operations, administration, and maintenance of cloud systems / virtualized servers, including 5+ years in a technical leadership role
- [Required] AWS Solutions Architect Associate Certification or higher strongly desired
- ...for personally driving reliability, observability, and cloud platform engineering excellence across a large, complex... ...who can personally raise the bar on SRE and observability practices, reduce... ....Reports to the Executive Director of Platform EngineeringBring deep,...SuggestedFull timeRemote work2 days per week
- ...Inc. (CCC) is a leading cloud platform for the multi-trillion-dollar... ...a highly skilled Platform Engineer with deep expertise in designing... ...:Enhance and evolve our observability capabilities across Azure, AWS... ...Site Reliability Engineering (SRE), Platform Engineering, or related...SuggestedFull time
$239k
...researchers on a secure, technology-enabled platform that drives clinical innovation and... ...Do? The Overview The Sr. Director, Platform Engineering & Tooling will be the senior leader... ...the foundational tooling (including observability, cloud infrastructure tooling,...SuggestedTemporary workFlexible hours$119.77k - $140.9k
...enterprise API ecosystem, leading engineering efforts across Apigee OPDK,... ...(Azure), and Apollo GraphQL platforms.Design and deliver secure,... ...through monitoring, observability, incident response, performance... ...Site Reliability Engineering (SRE).Proven ability to lead large...SuggestedFull timeWork experience placementLocal area3 days per week$120k - $140k
...the enterprise, and our cloud platforms are central to how our... ...Senior DevOps & Site Reliability Engineer, you will be hands-on at the... ...keeping those platforms reliable, observable, secure, and performant in... ...DevOps, cloud infrastructure, SRE, platform engineering, or related...SuggestedPermanent employmentTemporary workWork experience placementH1bLocal areaRemote workFlexible hours- ...DescriptionWe are seeking a Sr. DevOps Engineer to modernize and scale our... ...CI/CD, cloud infrastructure, platform engineering, automation, and... ..., Platform Engineering, SRE, Security, and Architecture... ...Integrate security, compliance, observability, and governance controls into...Remote work
$100.4k - $203k
...core values.ResponsibilitiesThe AWS Platform Engineering Manager is a hands-on technical leader... ...identity and access management, security, observability, and operational best practices.... ...unless we have an agreement signed by the Director of Talent Acquisition, SVP, to fill a...Temporary work$119.4k - $204.6k
...ResponsibilitiesDefine enterprise-wide platform strategy, vision, and target-state architectures... ...and guardrails across all platform engineering domains.Drive innovation in cloud, data,... ...engineering, cloud infrastructure, SRE roles or relatedIn Lieu of Education12 Years...Full timeTemporary workPart time$240k - $375k
...leading technology and exceptional service. The global lead for Platform Engineering leads the strategy, development, and operationalization of... ..., enablement, and measurement.Partner closely with Security, SRE, Risk, Control, and Architecture teams to embed best practices...H1bWorldwideFlexible hours$168.48k - $272.95k
...achieve. Together.SummaryThe Director, Engineering is responsible for leading... ...reusable engineering and platform capabilities that accelerate... ...including automation, CI/CD, observability, resiliency, testing,... ...infrastructure, operations, SRE, and end-user support teams...Full timeWork at officeRemote workRelocationVisa sponsorshipRelocation package$232k - $319k
...let's talk.The Infrastructure Platform and Shared Services TeamOkta... ...networking, K8s platform, Observability, automation platform & tooling... ...various initiatives across SRE & Infrastructure... ...velocity of SRE and product engineering by developing robust platforms...Permanent employmentLocal areaWorldwideFlexible hours$157.9k - $282.1k
Principal Full Stack Engineer, AI Platform & AgentsBuild the GenAI platform... ...developer tooling, CI/CD, and observability for safe, fast iteration (... ...’ll report directly to the Director of Engineering, AI Platform... ...Reliability Engineering (SRE)Quality engineering / testing...Full timeWork at officeRemote work2 days per week$131.75k - $170.5k
...this position is Senior Linux Engineer, this role has been posted... ...Overview We are seeking a Senior Platform Engineer to join our Systems... ...with Application Support (SRE), Development, and Security teams... ...management. Familiarity with observability tools such as Prometheus,...Full timeWork at officeImmediate start$106k - $117k
...ResponsibilitiesReporting to the Director of techstaff, acts as a... ...architectural standards, mentors junior engineers, drives platform reliability, automation, security, and observability across on-prem and cloud... ..., DevOps, platform, or SRE role supporting production systems...Full timeWork experience placement- ...RezinJob ID: REQ8338We are seeking a Trading Platform Engineer to join our Systematic Technology team, focused on the reliability, observability, and day-to-day support of a high-... ...QualificationsStrong experience in production engineering, SRE, or application support within real-time...
$87.5k - $145k
...a new commercial data platform serving institutional... ...loggingReliability & Observability: Build monitoring, alerting... ...: Work with the Data Engineer (who builds pipelines... ...and the Quantitative Director (whose outputs you deliver... ..., or DevOps/SRE — ideally supporting external...Local areaFlexible hours$165k - $225k
...Sr. Site Reliability Engineer (SRE) Chicago, IL or Remote Moonlite delivers high-performance... ...engineers, network engineers, and platform engineering team, you'll architect and... ...while establishing the automation, observability, and operational practices. Job...Remote workFlexible hours- ...them. We are seeking an AI Platform Engineer. Our Platform... ...spanning Delivery, Cloud, and SRE — is expanding to lead and maintain... ...monitoring, logging, and observability (tracking latency, quality,... ...SRE and reports to the Senior Director of Platform Engineering. We’...Full timeWork at officeRelocation
$211.5k - $235k
...care network with the most advanced care platform. Our August 2021 acquisition of Home... ...& Role We're looking for a hands-on Engineering Manager to lead Platform Engineering, a... ...level within a platform, infrastructure, or SRE organization at a product-centric company...Permanent employmentTemporary workWork at officeLocal areaRemote workRelocationHome officeFlexible hours$145k
...Platform Engineer at Akuna Capital Akuna Capital is an innovative trading firm with a strong... ...Platform Engineering, Cloud Engineering, SRE, DevOps, or related infrastructure... ...native platforms ~ Strong knowledge of observability, including monitoring, logging, alerting...InternshipWork at officeRemote work- ...Inc. (CCC) is a leading cloud platform for the multi-trillion-dollar... ...Lead Database Platform Engineer to drive the strategy, architecture... ...closely with Engineering, SRE, DevOps, Architecture, and Product... ..., Infrastructure as Code, observability platforms, and DevOps practices...Full time
$200k - $250k
...We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems... ...Build end-to-end performance profiling and observability to identify bottlenecks from individual... ..., and retirement. Partner with SRE and platform teams to automate model deployment...Temporary workFlexible hours$100k - $130k
...create your future.We are looking for a Platform Integration Engineer to join our engineering team and take... ...platform services — authentication, observability, configuration management, logging —... ..., or Pulumi Collaborate with DevOps/SRE to define deployment pipelines and reliability...Full timeWork experience placementLocal areaRemote workFlexible hours$212.5k - $275k
...market participants around the world. The goal of the Cboe Platform Engineering group focuses around building a scalable and secure foundations... ...) and on-premises Linux environments. Reporting to the Sr. Director, Platform Engineering, this role will lead a team of...Full time$225k - $280k
...’re looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward. ESM Platforms Engineering ManagerPurpose of the RoleJLL's Enterprise Service Management (ESM) team is hiring an ITSM Engineering Manager to lead the engineering...Full timeFor contractorsLocal areaShift work$130k - $190k
...Responsible for the operational reliability, observability, and stability of the Strategic Full Revaluation Capability (SFRC) batch platform. This role acts as the first line of... ..., and stability improvements. Champion SRE and DevOps best practices, including automation...Full timeTemporary workRemote workWorldwide$184k - $230k
...employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Overall Purpose The Principal Platform Security Engineer is a hands-on enterprise technical leader responsible for defining the long-term technical vision, secure target states,...Hourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$93k - $189k
DescriptionSenior Manager, AI Engineering PlatformsBuilding the Google-powered foundation... ...operation of Huntington's AI Engineering Platform on Google Cloud, delivering the shared... ...and state services, runtime execution, observability, platform governance, and engineering enablement...Full timeWork at officeRemote workWork from homeFlexible hours$204k - $306k
...The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes... ...Manager of Infrastructure Platform and Shared Services, you... ...networking, K8s platform, CI/CD, Observability, automation platform &... ...be doing Managing a team of SRE’s supporting various workloads...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director Platform Engineering - SRE / Observability. Be the first to apply!
- principal security engineer Chicago, IL
- chief engineer Chicago, IL
- senior chief engineer Chicago, IL
- principal infrastructure engineer Chicago, IL
- principal developer Chicago, IL
- general engineer Chicago, IL
- director software engineering Chicago, IL
- director of electrical engineering Chicago, IL
- engineering director Chicago, IL
- director data engineering Chicago, IL




