Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director of Platform Engineering and Observability

Request Technology, LLC

Job Description

***Position is bonus eligible***

\n

Prestigious Financial Institution is currently seeking a Director of Platform Engineering with strong SRE and Observability leadership experience. Candidate will lead and manage this team to drive reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing environment. The ideal candidate is a hands-on, data-driven technical leader, who can personally raise the bar on SRE and observability practices, reduce waste, optimize cloud efficiency, and improve performance, in close partnership with a dedicated SRE/monitoring team and a centralized architecture function.

\n

Responsibilities:

\n

Management Responsibilities:

\n
    \n
  • Manage and develop a team of engineers and managers focused on SRE, observability, and cloud platform engineering, growing and retaining talent
  • \n
  • Lead hiring and staffing for the team, including sourcing, interviewing, and selecting engineers and managers to build out the organization
  • \n
  • Meet with team members regularly to provide coaching and feedback on performance
  • \n
  • Perform evaluations and deal effectively with staff problems and corrective actions as needed
  • \n
  • Develop employee career development plans to assist with team member career growth and development
  • \n
  • Manage and participate in the implementation of production changes during defined maintenance windows and support on-call rotation
  • \n
  • Serve as a point of escalation within the team for reliability, observability, and platform support issues
  • \n
  • Foster an atmosphere of trust, respect, and high performance while displaying strong ethics and integrity
  • \n
  • Manage project and daily work task planning and prioritization, meeting project deadlines while maintaining a high quality of work
  • \n
  • Ensure team compliance with all appropriate policies and procedures, and institute corrective actions to address audit and other regulatory or compliance findings
  • \n
  • Operate within budget; establish and assure adherence to schedules, work plans, and performance requirements
  • \n
\n

Technical & Platform Responsibilities:

\n
    \n
  • Bring deep, hands-on SRE and observability expertise to Platform Engineering, elevating the maturity, rigor, and technical depth of the organization’s observability, monitoring, and reliability practices in close partnership with the dedicated SRE and monitoring team
  • \n
  • Serve as a technical authority on reliability engineering, mentoring engineers and managers and raising the bar on SLOs/SLIs, error budgets, incident response, and postmortem discipline across the platform organization
  • \n
  • Own the definition, governance, and lifecycle of SLOs and SLAs across Platform Engineering, codified as SLO-as-code (e.g., OpenSLO, Sloth, or Nobl9) and version-controlled in the repo rather than only living in a vendor UI
  • \n
  • Drive resilience engineering as a discipline across the platform, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing, to proactively validate and improve system resilience ahead of production incidents
  • \n
  • Deliver golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box, reducing toil and accelerating safe onboarding for engineering teams
  • \n
  • Act as product owner for observability within the platform organization, defining the roadmap and requirements for observability capabilities while the dedicated SRE and monitoring team retains ownership of the underlying tooling and operations
  • \n
  • Own and drive platform engineering delivery and reliability metrics, including DORA metrics (deployment frequency, lead time for changes, change failure rate, MTTR), cycle time, and throughput, holding the organization accountable to continuous improvement
  • \n
  • Define and report on a rounded set of platform and observability metrics: platform adoption (percentage of services on paved roads, service catalog completeness, time to first deploy, onboarding time), observability coverage (percentage of tier-1 services with SLOs, instrumented traces, and runbooks), observability health (MTTD, MTTA, alert actionability ratio, toil percentage, and paging load per engineer), observability economics (telemetry cost per service, cardinality and ingestion growth, and retention tiers), and developer experience (SPACE/DevEx measures alongside DORA)
  • \n
  • Champion the Golden Signals, RED, and USE methods as the platform’s common language for monitoring and alerting
  • \n
  • Ensure audit-grade telemetry practices, including log retention and immutability, PII and sensitive-data scrubbing, and access controls on observability data, mapped to relevant compliance frameworks (e.g., CIS, NIST, and SIFMU-specific resilience and reporting expectations)
  • \n
  • Drive reliability across four core pillars: a data-driven culture (“data junkies”) that uses telemetry and metrics, not intuition, to make decisions; waste reduction through elimination of toil and inefficient processes; cloud measures such as utilization, rightsizing, and FinOps metrics; and continuous performance optimization
  • \n
  • Translate reliability signals (SLOs, error budgets, incident learnings) surfaced by the SRE organization into concrete cloud platform investments and roadmap priorities
  • \n
  • Drive results in the cloud platform by building a reliable, scalable, secure technology stack in collaboration with engineers and leaders, using leading industry practices
  • \n
  • Champion infrastructure-as-code (IaC) practices across the cloud platform, ensuring provisioning, configuration, and environment management are automated, version-controlled, and consistently applied
  • \n
  • Assess and plan for capacity needs within the cloud platform and forecast accordingly
  • \n
  • Collaborate with the centralized architecture function on platform architecture decisions, ensuring cloud platform engineering execution aligns with broader architectural standards and direction
  • \n
  • Implement and manage initiatives within your assigned area of responsibility with accountability for results and compliance with all controls and security requirements
  • \n
  • Lead in the development of technology roadmaps and end-of-life technology plans
  • \n
  • Effectively communicate project and operational service issues to senior management promptly with observations, decisions, and recommendations for corrective measures
  • \n
  • Other duties as assigned
  • \n
  • Manage a team of engineers and engineering managers focused on SRE, observability, and cloud platform engineering. Full people-management responsibilities, including hiring, performance management, coaching, career development, and corrective action.
  • \n
\n

Qualifications:

\n
    \n
  • [Required] 5+ years of demonstrated experience leading engineering teams, with an emphasis on developing key talent and cultivating positive, high-performing cultures
  • \n
  • [Required] 10+ years of progressive, hands-on experience in software engineering with an understanding of large-scale computing solutions (primarily AWS), including software design and development, database architectures, IP networking, security, cloud operations, and performance tuning
  • \n
  • [Required] Demonstrated, hands-on expertise building and maturing SRE and observability practices at scale, able to personally raise the technical bar for a dedicated SRE/monitoring team, not just consume their output
  • \n
  • [Required] Demonstrated track record owning the definition and governance of SLOs and SLAs for mission-critical systems, and personally driving resilience engineering practices (chaos engineering, load/performance testing) to validate reliability ahead of incidents
  • \n
  • [Required] Demonstrated track record driving reliability through a data-driven culture, waste/toil reduction, cloud efficiency measures, and performance optimization, treating metrics as the primary lens for every decision
  • \n
  • [Required] Experience defining, instrumenting, and acting on software delivery performance metrics (DORA metrics, cycle time, deployment frequency, lead time, etc.) to drive engineering improvement initiatives
  • \n
  • [Required] Strong consultative, communication, team player, and analytical skills, with the ability to regularly interact between various teams distributed across the US
  • \n
  • [Required] Strong technical team leadership and technical project management skills
  • \n
  • [Required] Relevant experience leading highly technical team members through adopting new technologies while maintaining highly available, mission-critical systems, with a proven track record of success
  • \n
  • [Required] Ability to clearly communicate verbally and in writing to business and technology leaders, architects, developers, and team members
  • \n
  • [Required] Must be able to collaborate effectively with a group of high-performing, technical individuals
  • \n
  • [Required] Experience acting as a product owner, defining roadmap, requirements, and priorities, for a platform capability such as observability, ideally in partnership with a separate team that owns the underlying tooling and operations
  • \n
  • [Required] Experience with architecting, implementing, and maintaining highly available mission-critical environments for 24x7 availability
  • \n
  • [Required] Demonstrated history of working within deadlines and ability to work well under pressure
  • \n
  • [Required] Experience managing work tasks using Agile methodology/scrum desired
  • \n
  • [Preferred] Comfort with ambiguity and demonstrated ability to lead complex programs in a decentralized environment
  • \n
  • [Preferred] Experience working in an environment with a defined production change control process; experience working with audits and compliance or in a regulated environment a plus
  • \n
  • [Preferred] Experience in organizations with a mature, centralized SRE function, avoiding role/scope overlap
  • \n
  • [Required] Deep expertise in OpenTelemetry, including instrumentation standards, auto-instrumentation, semantic conventions, and the OTel Collector, as the foundation for a vendor-neutral, paved-road instrumentation strategy
  • \n
  • [Required] Hands-on experience with metrics engines and time-series databases, including Prometheus and PromQL, plus scale-out options such as Mimir, Thanos, VictoriaMetrics, or Amazon Managed Prometheus, including cardinality management
  • \n
  • [Required] Hands-on experience with tracing and logging backends such as Tempo/Jaeger, Loki/Elastic/Splunk (Splunk is common in financial services), and AWS X-Ray
  • \n
  • [Required] Experience with Kubernetes and Kafka observability specifically, including EKS metrics, kube-state-metrics, consumer lag, and broker health
  • \n
  • [Required] Experience with alerting and incident tooling such as PagerDuty or Opsgenie, including ServiceNow integration, alert routing, and noise reduction
  • \n
  • [Required] Hands-on experience with SLO-as-code frameworks such as OpenSLO, Sloth, or Nobl9, defining and governing SLOs in the repo rather than only in a vendor UI
  • \n
  • [Required] Experience delivering golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box
  • \n
  • [Required] Experience with resilience validation practices, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing
  • \n
  • [Required] Deep, hands-on mastery of observability tooling and practices (metrics, distributed tracing, centralized logging, dashboards, alerting), e.g., Datadog, Prometheus/Grafana, CloudWatch, or equivalent, sufficient to elevate, not just consume, a dedicated SRE team’s capability
  • \n
  • [Required] Deep understanding of SRE principles including SLOs/SLIs, error budgets, incident management, and postmortem culture, with a track record of driving adoption and maturity
  • \n
  • [Required] Fluency in the Golden Signals, RED, and USE methods as applied frameworks for monitoring and alerting design
  • \n
  • [Required] Hands-on experience with: Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, Github, and configuration management tools such as Puppet, Chef, or Ansible
  • \n
  • [Required] Deep, hands-on expertise with infrastructure-as-code (IaC) tools and practices (e.g., Terraform, CloudFormation, CDK, Pulumi), with a track record of driving IaC adoption at scale across a cloud platform organization
  • \n
  • [Required] Relevant experience with configuration and implementation of IaaS, Infrastructure as Code, AWS, Azure, etc.
  • \n
  • [Required] Expert working knowledge of infrastructure design and components, such as servers, operating systems, networks, and storage
  • \n
  • [Required] Basic understanding of good delivery practices and continual integration and improvement; Agile/Lean background for projects and project delivery
  • \n
  • [Preferred] Experience with telemetry pipeline tools such as Fluent Bit, Vector, or Cribl for routing, sampling, redaction, and cost control
  • \n
  • [Preferred] Experience with synthetic monitoring, real user monitoring (RUM), and eBPF-based observability
  • \n
  • [Preferred] Familiarity with audit-grade telemetry practices (log retention/immutability, PII and sensitive-data scrubbing, access controls on observability data) and mapping observability controls to compliance frameworks such as CIS, NIST, or SIFMU-specific resilience requirements
  • \n
  • [Preferred] Familiarity with engineering metrics/analytics platforms (e.g., LinearB, Jellyfish, Sleuth, Haystack, or internal equivalents) used to track DORA metrics and delivery performance
  • \n
  • [Preferred] Competent in all phases of application development and implementation, including SDLC; hands-on scripting/development skills in Python, Ruby, Go, Java, etc. in a corporate environment strongly desired
  • \n
  • [Preferred] Experience establishing IaC governance and standards (module libraries, policy-as-code, drift detection) across multiple teams or business units
  • \n
  • [Preferred] Experience building a metrics-driven engineering culture, including scorecards, dashboards, or leadership reporting on delivery performance
  • \n
  • [Required] Bachelor’s degree, preferably in a technical discipline (Computer Science, Mathematics, etc.), or equivalent combination of education and experience required; Master’s degree and relevant experience also considered
  • \n
  • [Required] 10+ years’ experience in IT systems installation, operations, administration, and maintenance of cloud systems / virtualized servers, including 5+ years in a technical leadership
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Director of Platform Engineering and Observability in Chicago, IL vacancy
  •  ...Director of Cloud Platform Engineering (SRE & Observability) Job description: This is a people-management role with full supervisory responsibility for a team of engineers and managers. The Director will lead and manage this team to drive reliability, observability... 
    Suggested
    Full time

    New York Technology Partners

    Chicago, IL
    1 day ago
  • $175k - $240k

     ...Director, Platform Engineering SALARY: $175k - $240k plus 30% bonus LOCATION: CHICAGO, IL HYBRID 3 DAYS ONSITE You will manage...  ...of SREs and managers. 7-10 people and managers. Keys SRE observability infrastructure kubernetes kafka aws terraform. j The... 
    Suggested
    Full time

    Request Technology, LLC

    Chicago, IL
    2 days ago
  •  ...Job Description Director, Platform Engineering - Financial Services \n Location: Chicago, IL \n \n We are seeking a Director of Platform...  ...7–10 engineers and managers focused on SRE, observability, reliability, and cloud platform engineering . This leader... 
    Suggested

    Request Technology, LLC

    Chicago, IL
    2 days ago
  • $100.4k - $203k

     ...core values.ResponsibilitiesThe AWS Platform Engineering Manager is a hands-on technical leader...  ...identity and access management, security, observability, and operational best practices....  ...unless we have an agreement signed by the Director of Talent Acquisition, SVP, to fill a... 
    Suggested
    Temporary work

    Old National Bank

    Chicago, IL
    4 days ago
  •  ...role responsible for personally driving reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing...  ...primary duty satisfactorily.Reports to the Executive Director of Platform EngineeringBring deep, hands-on SRE and... 
    Suggested
    Full time
    Remote work
    2 days per week

    The Options Clearing Corporation

    Chicago, IL
    4 days ago
  •  ...Solutions Inc. (CCC) is a leading cloud platform for the multi-trillion-dollar insurance...  ...are seeking a highly skilled Platform Engineer with deep expertise in designing, deploying...  ...:Enhance and evolve our observability capabilities across Azure, AWS, and application... 
    Full time

    CCC Information Services

    Chicago, IL
    4 days ago
  • $239k

     ...researchers on a secure, technology-enabled platform that drives clinical innovation and...  ...Do? The Overview The Sr. Director, Platform Engineering & Tooling will be the senior leader...  ...the foundational tooling (including observability, cloud infrastructure tooling,... 
    Temporary work
    Flexible hours

    Prolaio

    Chicago, IL
    3 days ago
  • $119.77k - $140.9k

     ...for our enterprise API ecosystem, leading engineering efforts across Apigee OPDK, Apigee Hybrid (Azure), and Apollo GraphQL platforms.Design and deliver secure, scalable, and...  ...platform reliability through monitoring, observability, incident response, performance tuning, and... 
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Chicago, IL
    1 day ago
  •  ...DescriptionWe are seeking a Sr. DevOps Engineer to modernize and scale our software delivery...  ...on CI/CD, cloud infrastructure, platform engineering, automation, and developer...  ...toil. Integrate security, compliance, observability, and governance controls into delivery... 
    Remote work

    Inspira Financial

    Oak Brook, IL
    2 days ago
  • $135.2k - $236.6k

    Director of Engineering - Data Platforms page is loaded## Director of Engineering - Data Platformslocations: Edina, MN 55435: Chicago, IL 60607time type...  ...teams to deliver reliable, secure, scalable, observable, and high-performing enterprise data services.* Drive... 

    Jobleads-US

    Chicago, IL
    1 day ago
  • $212.5k - $275k

     ...market participants around the world.  The goal of the Cboe Platform Engineering group focuses around building a scalable and secure foundations...  ...) and on-premises Linux environments. Reporting to the Sr. Director, Platform Engineering, this role will lead a team of... 
    Full time

    Cboe Exchange

    Chicago, IL
    3 days ago
  • $225k - $280k

     ...’re looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward. ESM Platforms Engineering ManagerPurpose of the RoleJLL's Enterprise Service Management (ESM) team is hiring an ITSM Engineering Manager to lead the engineering... 
    Full time
    For contractors
    Local area
    Shift work

    Jones Lang LaSalle

    Chicago, IL
    3 days ago
  • $240k - $375k

     ...partners, we serve the world’s most sophisticated clients using leading technology and exceptional service. The global lead for Platform Engineering leads the strategy, development, and operationalization of shared developer tools and foundational services used across... 
    H1b
    Worldwide
    Flexible hours

    Northern Trust

    Chicago, IL
    4 days ago
  • $119.4k - $204.6k

    Essential ResponsibilitiesDefine enterprise-wide platform strategy, vision, and target-state architectures for platforms such as Microsoft...  ....Establish standards and guardrails across all platform engineering domains.Drive innovation in cloud, data, and AI platforms (Fabric... 
    Full time
    Temporary work
    Part time

    Alliant Credit Union

    Chicago, IL
    15 hours ago
  •  ...Vizient, Inc. is seeking a Director of Engineering for Data Platforms to define and execute the enterprise data platform strategy, roadmaps, and modernization efforts across Azure Databricks and Lakehouse architectures. You will lead platform engineering and administration... 

    Jobleads-US

    Chicago, IL
    1 day ago
  • $184k - $230k

     ...employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Overall Purpose The Principal Platform Security Engineer is a hands-on enterprise technical leader responsible for defining the long-term technical vision, secure target states,... 
    Hourly pay
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning

    Chicago, IL
    15 hours ago
  • $232k - $319k

     ...are too, let's talk.The Infrastructure Platform and Shared Services TeamOkta authenticates...  ...on Edge networking, K8s platform, Observability, automation platform & tooling. What you...  ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms,... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Chicago, IL
    2 days ago
  • $131.75k - $170.5k

     ...title for this position is Senior Linux Engineer, this role has been posted externally...  ....Role Overview We are seeking a Senior Platform Engineer to join our Systems Platform Engineer...  ...management. Familiarity with observability tools such as Prometheus, Grafana, or Loki... 
    Full time
    Work at office
    Immediate start

    Cboe Exchange

    Chicago, IL
    3 days ago
  • $157.9k - $282.1k

    Principal Full Stack Engineer, AI Platform & AgentsBuild the GenAI platform that powers critical...  ....Build developer tooling, CI/CD, and observability for safe, fast iteration (evals, canaries...  ...title. You’ll report directly to the Director of Engineering, AI Platform.Team Size... 
    Full time
    Work at office
    Remote work
    2 days per week

    Wolters Kluwer

    Chicago, IL
    3 days ago
  • $87.5k - $145k

     ...layer for a new commercial data platform serving institutional...  ...audit loggingReliability & Observability: Build monitoring, alerting,...  ...enforcementCollaboration: Work with the Data Engineer (who builds pipelines that...  ...), and the Quantitative Director (whose outputs you deliver... 
    Local area
    Flexible hours

    alterDomus

    Chicago, IL
    1 day ago
  •  ...Experience ProfessionalsContact: Ashley RezinJob ID: REQ8338We are seeking a Trading Platform Engineer to join our Systematic Technology team, focused on the reliability, observability, and day-to-day support of a high-performance, low-latency trading platform.This role... 

    Balyasny Asset Management

    Chicago, IL
    3 days ago
  • $106k - $117k

     ...production.ResponsibilitiesReporting to the Director of techstaff, acts as a senior technical leader...  ...architectural standards, mentors junior engineers, drives platform reliability, automation, security, and observability across on-prem and cloud environments.Lead initiatives... 
    Full time
    Work experience placement

    The University of Chicago

    Chicago, IL
    3 days ago
  • $130k - $150k

     ...cloud technologies is essential for this role. The Senior Platform Engineer is responsible for designing, building, and continuously improving...  ...strategy through platform standardization, governance, observability, cost optimization, and automation-first operational... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Chicago, IL
    1 day ago
  •  ...conception to deployment. As a Senior Software Engineer, you will have relevant experience with...  ...patterns, Cloud foundational patterns, Observability patterns, Developer experience patterns...  ...with one or more cloud platforms, preferably GCPExcellent communication... 

    Inspira Financial

    Oak Brook, IL
    1 day ago
  • $135.2k - $236.6k

     ...strategic and technical leadership across engineering teams delivering AI-enabled intelligent...  ...as Code, automated testing, and observability.Lead the delivery of secure, scalable,...  ...Microsoft Azure or other public cloud platforms required.Strong understanding of modern... 

    Vizient

    Chicago, IL
    4 days ago
  •  ...unlock incredible career growth opportunities, join us, and build real world value. Join Ripple Labs Inc. as the Manager, AI Platform Engineering in Chicago, IL, and spearhead our bold venture in AI platform development! This remarkable opportunity enables you to craft... 
    Full time
    Work at office
    Local area

    Hidden Road

    Chicago, IL
    3 days ago
  • Job DescriptionWe are seeking a Lead Platform Software Engineer to join our growing team. This role is responsible for the full software development...  ...Delivery patterns, Cloud foundational patterns, Observability patterns, Developer experience patterns. Strong analytical... 
    Temporary work
    Remote work

    Inspira Financial

    Oak Brook, IL
    2 days ago
  •  ...RezinJob ID: REQ8278About UsOur Database Engineering team is transforming how Balyasny Asset...  ...services through a standardized platform built on declarative, code-defined provisioning...  ....· Develop GitOps workflows and observability pipelines for comprehensive monitoring... 

    Balyasny Asset Management

    Chicago, IL
    2 days ago
  • $211.5k - $235k

     ...Platform Engineering Manager - Infra + DevOps Remote Position Honor Technology’s mission is to change the way society cares for older adults. As a leader in aging care innovation, Honor provides the technology, tools, and services that empower older adults to live... 
    Permanent employment
    Temporary work
    Work at office
    Local area
    Remote work
    Relocation
    Home office

    Jobleads-US

    Chicago, IL
    2 days ago
  •  ...Honor is hiring a hands-on Platform Engineering Manager to lead Infra + DevOps for their Care Platform. This remote-first role guides a senior team, owns the platform roadmap, and partners with ICs and domain teams to deliver scalable cloud infrastructure and developer... 
    Remote job

    Jobleads-US

    Chicago, IL
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director of Platform Engineering and Observability. Be the first to apply!