Senior Lead Site Reliability Engineer
J.P. Morgan
hackajob is collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms and Foundational Services (IPFS) team, you work with your fellow stakeholders to define non-functional requirements (NFRs) and availability targets for the services in your application and product lines. You will ensure those NFRs are accounted for in your products' design and test phases, that your service level indicators are effectively measuring customer experience, and that service level objectives are defined with stakeholders and implemented in production. Job Responsibilities Creates and delivers high quality designs, roadmaps, and program charters alongside the engineering team Acts as a key resource and mentor for technologists in your area seeking advice on technical and business issues, and serves as a culture carrier and site reliability adoption champion for your team Collaborates with others to create and implement observability and reliability designs for complex systems which are robust, stable, and do not incur additional toil or technical debt Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements. Drives evolution and debugging of critical components by understanding application and platform interdependencies and limitations Provides comprehensive and ongoing guidance, tools, and solutions to support the firms' growth Make significant contributions to JPMorganChase's site reliability community via internal forums, communities of practice, guilds, and conferences Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls. Required qualifications, capabilities, and skills Formal training or certification on site reliability engineering concepts and 5 years applied experience Advanced knowledge in site reliability culture and principles with demonstrated ability to implement site reliability within an application or platform Advanced knowledge and experience in observability such as white and black box monitoring, service level objectives, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Expert-level proficiency in Java, Go (Golang), Python, and Terraform for building enterprise-grade applications, high-performance systems, automation, and infrastructure as code Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity. Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes. Advanced knowledge of software applications and technical processes with considerable depth in multiple technical disciplines including distributed systems, microservices architecture, and cloud-native technologies Hands-on experience building AI Agents and autonomous systems with proficiency in AI frameworks (LangChain, LangGraph, AutoGen, CrewAI) and leveraging AI development tools (GitHub Copilot, Claude, etc.) to accelerate development and innovation and Expertise in designing and implementing logging pipelines (Fluentd, Logstash, Vector) and systems for metrics collection, analysis, and distributed tracing Strong experience building production-grade RESTful APIs and designing message queue architectures (Kafka, RabbitMQ, SQS) for event-driven systems; and expertise in graph databases (Neo4j, TigerGraph), vector databases (Pinecone, Weaviate, Chroma), and integrating multiple data stores for AI-powered systems Proficiency with containerization (Docker, Kubernetes), CI/CD pipelines, and GitOps workflows Ability to communicate data-based solutions with complex reporting and visualization methods, recognized as an active contributor of the engineering community, and continues to expand network and leads evaluation sessions with vendors to see how offerings can fit into the firm's strategy Preferred qualifications, capabilities, and skills Experience with MCP (Model Context Protocol) Servers or similar agent frameworks for building autonomous systems, and understanding of LLM integration, prompt engineering, and RAG (Retrieval-Augmented Generation) Familiarity with AI/ML model building, deployment, and lifecycle management using frameworks like TensorFlow, PyTorch, or scikit-learn Experience with big data technologies (Hadoop, Spark, Flink), analytical databases, NoSQL databases (MongoDB, Cassandra, DynamoDB), and time-series databases (InfluxDB, TimescaleDB) Knowledge of security best practices and compliance requirements in highly regulated industries, with experience in chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos) and GameDay exercises Contributions to open-source projects, particularly in SRE, observability, or AI/ML domains, and certifications in cloud platforms (AWS, Azure, GCP) Strong communication skills with ability to mentor and educate others on site reliability principles and practices, and ability to anticipate, identify, and troubleshoot defects found during testing This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase's review of criminal conviction history, including pretrial diversions or program entries. ABOUT US JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world's most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation. JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans ABOUT THE TEAM Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we're setting our businesses, clients, customers and employees up for success.aa415a4b-8b21-40fc-a65c-70d2b25ca29a
- ...Senior Lead Site Reliability Engineer Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan...Senior
$90k - $180k
...changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices,... .... JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale...SeniorFull timeRemote workShift work- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...Senior
$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...Senior$174k - $253k
...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent... ...distributed systems. 2 years of experience leading projects and providing technical leadership... ...or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when...Senior- ...Ernst & Young Oman is looking for a Real Estate Tax Senior Manager to lead tax planning projects and provide advisory services across various sectors. This role requires strong analytical skills and a proactive attitude to improve client tax activities. The ideal candidate...Senior
- Recruiter / Talent Acquisition Junior / Senior / Lead About Us Catalyst Labs is a specialized talent agency with a specialized vertical... ...organizations. Our clients are building high-performance teams across engineering, product, finance, GTM, and operations and they rely on...SeniorRemote work
- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer... ...operational maturity in PCI-DSS-regulated environments. Lead incident management during SEV1/SEV2 events and improve response...Senior
- ..., and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...SeniorRemote work
- A local youth sports organization in Mountain View seeks a Sports Coordinator responsible for enhancing the sports experience for players and coaches. This role involves empowering volunteer coaches, supervising game day operations, and teaching core values while maintaining...SeniorHourly payLocal areaWeekend workAfternoon shift
- Citi is seeking a Private Banker Sr. Principal in Palo Alto, California to effectively manage high-net-worth client relationships and serve as a trusted financial advisor. This role requires a broad understanding of financial strategies and the ability to identify client...Senior
$148k - $222k
...small, highly motivated, and focused on engineering excellence. This organization is for individuals... ...scalable, secure SaaS integrations. This senior individual contributor role serves as a... ...growth. RESPONSIBILITIES: Lead the design, implementation, administration...SeniorPermanent employmentFull timeTemporary work$151.6k - $245.3k
...Site Reliability Engineer Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability... ...with SRE and Dev teams in the on-call rotation Lead root cause analysis of critical business and production issues...$160k - $190k
...Applied Intuition is seeking a Risk and Compliance Lead in Sunnyvale, CA. This role involves leading security compliance initiatives and managing the security GRC program. The ideal candidate has over 6 years of experience in risk management and must be comfortable presenting...Senior- ...Lenmar Consulting Inc in Menlo Park, CA is seeking an Executive Assistant to support senior professionals in a fast-paced, client‑facing environment. The role requires strong judgment, attention to detail, and the ability to manage competing priorities with minimal direction...Senior
- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design... ...any immigration-related benefits. About USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security...
- ...together safely, transparently, and at scale. Join Wand in leading the Agentic Shift Wand is building a high-performing... ...a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep...Shift work
$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time- ...ActiveHours is looking for an experienced DevOps Engineer to enhance platform automation and site reliability in a collaborative environment. The role involves automating key systems, improving visibility through metrics, and troubleshooting critical problems. You'll...
- ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments...
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Work at office
$217.57k - $260k
...explicitly states otherwise, all roles are on-site five days per week at one of our... ...Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high... ...experience operating at this scale and leading infrastructure through significant transformation...Full timeTemporary workWork at officeRemote workFlexible hoursShift work- ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within...
$160.36k - $240.54k
...more connected future. About the Role We’re looking for senior engineers to build/scale Nuro's large-scale computing infrastructure in... ..., you’ll be working on building a scalable, efficient and reliable system that bridges the gap between fundamental infrastructure...SeniorFull time$220k - $255k
...A leading electric mobility company in Palo Alto is seeking a Reliability Engineer to drive the reliability of electric mobility vehicles during their development cycle. The ideal candidate will have experience in reliability engineering, strong technical knowledge of...SeniorFlexible hours$162k - $260k
...the location and motion of vehicles traveling on highways, surface streets, and endpoint facilities. We are searching for a Senior Software Engineer to join us in solving these technical challenges. In this role, you will Design, implement, and maintain solutions...SeniorFull time$204k - $259k
...geared towards both scaling models with efficiency and solving problems unique to ML for autonomous driving. We are looking for engineers with ML system expertise to help us train and improve pre-trained models to be deployed into Waymo Driver, and potential future...SeniorFull timeRemote work$152k - $214k
...Design APIs and SDKs together with external teams Support engineering development lifecycle processes in a highly regulated and safety... ...requirements and to test new ideas Mentor other engineers Lead technical discussions, feature development, and architecture reviews...SeniorFull time$235.03k - $352.29k
...Nuro’s autonomy stack utilizes an industry-leading sensor suite. Our tools must handle the... ...all of our users. The system must be reliable and scalable. This includes everything from... ...team closely collaborates with autonomy engineers to ensure our labeled data is high-...SeniorFull time- ...Senior Team Leader Crisis24, a GardaWorld company, is widely regarded as the leading integrated risk management, crisis response, consulting, and global protective solutions... ...of the client as the trusted, senior most on-site leader. Scheduling, personnel management,...SeniorLocal areaShift workNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Palo Alto, CA
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- senior technical product manager Palo Alto, CA
- senior medical science liaison Palo Alto, CA
- senior accountant remote Palo Alto, CA
- senior marketing account manager Palo Alto, CA
- senior robotics software engineer Palo Alto, CA
- senior dynamics crm developer Palo Alto, CA
- senior compensation manager Palo Alto, CA



