Lead Site Reliability Engineer
Hackajob
Lead Site Reliability Engineer
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
- Supports the adoption of site reliability engineering best practices within your team
- Production 24*7 support for business-critical applications
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
- Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
- Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
- Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
- Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
- Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
Preferred qualifications, capabilities, and skills
- Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
- Proven track record of initiating and executing ideas that address complex business challenges
- Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
- Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
- Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
- Experience in handling critical incident and change management – be part of critical incident taskforce call.
- Familiarity of agile practices – preferably, scrum and Kanban
- ...ServicesSelling Points Contribute to the reliability of a high-transaction payment... ...principles.Job DescriptionSite Reliability Engineer OverviewThe Site Reliability Engineer ensures the... ...root cause analysis for incidents and lead post-mortems to capture lessons learned...SuggestedRemote work
$121.4k - $218.6k
...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing... ...building technical runbooks, leading complex incident response bridges, and...SuggestedWork experience placementWork at office- Job Title Primary Skill PCF (Pivotal Cloud Foundry) and Mongo DB Exposure to at least 1 Observability Tool such as AppDynamics, Splunk, Grafana Change Mgmt using CI/CD pipeline. Harness or equivalent tools Secondary Skill SSL Certificate management...Suggested
$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...SuggestedRemote work$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$87.1k - $157.45k
...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across... ...to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying...Local areaImmediate startWork from homeFlexible hours$75.7k - $136.3k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company...Work experience placementWork at office- ...Life Cycle (SDLC) to enhance automation, reliability, and delivery efficiency. Apply... ...driven decision-making. Learn and apply engineering processes, methodologies, and best... ...certification in Software Engineering or Site Reliability Engineering. ~6+ years of...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a...
$169.3k - $304.7k
...maintaining fast, efficient, scalable, and reliable routing software and infrastructure... ...global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... ...Plan (ESPP). Akamai provides industry‑leading benefits including healthcare, 401K savings...Work experience placementWork at office$84.9k - $209.5k
...Help ensure healthcare professionals can reliably access the applications they depend on... ...Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability,... ...provisioning health, and capacity. Lead complex troubleshooting. Coordinate technical...Temporary workImmediate startFlexible hoursShift work$91.2k - $136.8k
...Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an insurance... ...position will play a crucial role to lead infrastructure resilience in ensuring... ...experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or...Full timeTemporary workWork at office3 days per week- DescriptionJob Description SummaryThe Digital Site Reliability Engineer (SRE) - GCP Cloud Adoption Engineer is responsible for facilitating the... ...delivery of cloud-based applications.Incident Management: Lead incident response for cloud-related issues, conduct root cause...Full timeH1bWork at officeRemote workWork from homeFlexible hours
- ...company dedicated to making a positive impact on people's lives. Position Summary Reporting to the Director of Engineering, the Lead Site Reliability Engineer (SRE) is a senior technical contributor responsible for building reliable, scalable software systems and...Full timeRemote workMonday to FridayShift workNight shiftWeekend work
$61k - $101k
...Requirements: We require formal training or certification in site reliability engineering concepts, along with 5+ years of applied experience. We... ...reliability, performance, security, and cost, while leading on-call and major incidents, driving toil reduction, and applying...Full time$61k - $101k
...Requirements: We require formal training or certification in site reliability engineering concepts, along with 5+ years of applied experience. We... ...through internal forums and communities of practice. We lead initiatives that improve the reliability and stability of...Full time- ...platforms, applying strong experience in Ansible, continuous integration and continuous delivery practices, DevOps and Site Reliability Engineering to design, automate, and optimize geospatial data services.Partner with cross functional teams to ensure reliable map based...
- ...design effective solutions. Responsibilities include designing, testing, debugging, documenting, and supporting production systems; leading projects; mentoring teammates; and balancing development with support in a fast-paced, multi-office environment. #J-18808-Ljbffr...Work at office
- ...automated testing, quality assurance, and compliance. Overall expectations include demonstrated supervisory skills and ability to lead a team, effective organizational, communication and technology skills along with discipline-specific technical skills. The leader should...Flexible hours
- PrimeFlight Aviation Services, Inc. is seeking a Ramp Supervisor to lead a team of ramp agents and ensure safe, efficient ground handling of aircraft. You will oversee baggage handling, aircraft towing, and servicing, coordinating with flight crews and airport operations...
- ...department. The role provides supervision of department operations, directs human resources, and ensures regulatory compliance in a high reliability organization. You will oversee workflow, training, and scheduling; participate in process improvement, investigate specimen...
- ...accurate and timely while coordinating with multiple departments and payer workflows to support efficient patient access. You will lead staff, conduct formal performance reviews, enforce policies, manage staffing levels, and resolve billing and access issues, all while...
- A nationwide construction firm is seeking a Shelving & Racking Supervisor. The role involves installing and assembling steel racking systems and requires mechanical skills and ability to operate forklifts. Candidates must be comfortable working at heights and able to travel...
- OhioHealth Riverside Methodist Hospital in Columbus, OH is seeking a Histology ASCP Supervisor to oversee the laboratory's surgical pathology section. The role requires supervising technical operations, staff scheduling, training, and ensuring compliance with CAP, CLIA,...Night shift
- Ability to facilitate and lead discussions with business SME's • Ability to create and present Ai solution proposals and communicate both business and technical concepts. • Ability to lead and conduct emerging technology analysis and conduct proof of concept initiatives...
- ...scalable security, compliant digital transformation, and service optimization for a high-profile public sector client. The candidate will lead complex development, integrations, and governance efforts onsite in Columbus, OH, leveraging advanced JavaScript, HTML/CSS, and...
- JPMorganChase is seeking a Technology Support III to join the Mainframe and Mid-Range Compute Site Reliability and Engineering team. You will support end-to-end infrastructure services, collaborate with stakeholders, and ensure high availability across a globally distributed...
- G2O, based in Columbus, Ohio, is seeking a Guidewire Developer to design and implement integration solutions. Ideal candidates will have 3-5 years of Guidewire development experience and a Guidewire ClaimCenter certification. This role involves collaborating with a diverse...
$97k - $180k
...close briefings, providing Controllership leaders with visibility into key topics. This position offers significant opportunity to lead process improvements and advance technology-enabled efficiencies. The ideal candidate brings strong analytical and problem-solving skills...Full timeTemporary workPart timeCasual workWork at office- ...clients shift to the New using leading-edge technologies on some of... ...ensuring quality, value, and reliability of deployed systems.The Work:... ...business problems at client sites• Work with large-scale datasets... ..., business owners, engineers, architects, and UI designers...Full timeWork experience placementLive inWork at officeLocal areaShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead infrastructure engineer Columbus, OH
- lead security engineer Columbus, OH
- lead engineer Columbus, OH
- lead operating engineer Columbus, OH
- lead system engineer Columbus, OH
- lead web developer Columbus, OH
- lead network engineer Columbus, OH
- site reliability engineer sre Columbus, OH
- site reliability engineer Columbus, OH
- remote website tester Columbus, OH


