Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Lead Site Reliability Engineer

Chase

Senior Lead Site Reliability Engineer

Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms and Foundational Services (IPFS) team, you work with your fellow stakeholders to define non-functional requirements (NFRs) and availability targets for the services in your application and product lines. You will ensure those NFRs are accounted for in your products' design and test phases, that your service level indicators are effectively measuring customer experience, and that service level objectives are defined with stakeholders and implemented in production.

Job Responsibilities

  • Creates and delivers high quality designs, roadmaps, and program charters alongside the engineering team
  • Acts as a key resource and mentor for technologists in your area seeking advice on technical and business issues, and serves as a culture carrier and site reliability adoption champion for your team
  • Collaborates with others to create and implement observability and reliability designs for complex systems which are robust, stable, and do not incur additional toil or technical debt
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
  • Drives evolution and debugging of critical components by understanding application and platform interdependencies and limitations
  • Provides comprehensive and ongoing guidance, tools, and solutions to support the firm's growth
  • Make significant contributions to JPMorganChase's site reliability community via internal forums, communities of practice, guilds, and conferences
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Advanced knowledge in site reliability culture and principles with demonstrated ability to implement site reliability within an application or platform
  • Advanced knowledge and experience in observability such as white and black box monitoring, service level objectives, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
  • Expert-level proficiency in Java, Go (Golang), Python, and Terraform for building enterprise-grade applications, high-performance systems, automation, and infrastructure as code
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
  • Advanced knowledge of software applications and technical processes with considerable depth in multiple technical disciplines including distributed systems, microservices architecture, and cloud-native technologies
  • Hands-on experience building AI Agents and autonomous systems with proficiency in AI frameworks (LangChain, LangGraph, AutoGen, CrewAI) and leveraging AI development tools (GitHub Copilot, Claude, etc.) to accelerate development and innovation and Expertise in designing and implementing logging pipelines (Fluentd, Logstash, Vector) and systems for metrics collection, analysis, and distributed tracing
  • Strong experience building production-grade RESTful APIs and designing message queue architectures (Kafka, RabbitMQ, SQS) for event-driven systems; and expertise in graph databases (Neo4j, TigerGraph), vector databases (Pinecone, Weaviate, Chroma), and integrating multiple data stores for AI-powered systems
  • Proficiency with containerization (Docker, Kubernetes), CI/CD pipelines, and GitOps workflows
  • Ability to communicate data-based solutions with complex reporting and visualization methods, recognized as an active contributor of the engineering community, and continues to expand network and leads evaluation sessions with vendors to see how offerings can fit into the firm's strategy

Preferred qualifications, capabilities, and skills

  • Experience with MCP (Model Context Protocol) Servers or similar agent frameworks for building autonomous systems, and understanding of LLM integration, prompt engineering, and RAG (Retrieval-Augmented Generation)
  • Familiarity with AI/ML model building, deployment, and lifecycle management using frameworks like TensorFlow, PyTorch, or scikit-learn
  • Experience with big data technologies (Hadoop, Spark, Flink), analytical databases, NoSQL databases (MongoDB, Cassandra, DynamoDB), and time-series databases (InfluxDB, TimescaleDB)
  • Knowledge of security best practices and compliance requirements in highly regulated industries, with experience in chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos) and GameDay exercises
  • Contributions to open-source projects, particularly in SRE, observability, or AI/ML domains, and certifications in cloud platforms (AWS, Azure, GCP)
  • Strong communication skills with ability to mentor and educate others on site reliability principles and practices, and ability to anticipate, identify, and troubleshoot defects found during testing

This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase's review of criminal conviction history, including pretrial diversions or program entries.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Lead Site Reliability Engineer in Palo Alto, CA vacancy
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems...  ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    4 days ago
  •  ...Chase & Co. Payments is seeking a Vice President, Applied AI/ML Lead to own end-to-end delivery of high-impact AI/ML capabilities...  ...lead reviews, and mentor teams while partnering with Product, Engineering, Risk, Compliance, and Data teams #J-18808-Ljbffr Jobleads-US
    Senior

    Jobleads-US

    Palo Alto, CA
    6 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide...  ...dependencies, and shared infrastructure components. Lead or support incident triage for service degradation involving... 
    Senior
    Contract work

    VDart

    Santa Clara, CA
    3 days ago
  •  ...Team: Infra Reliability • SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma...  ...prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference...  ...metal role for a first-principles Linux engineer. You'll be the final escalation for the... 
    Senior
    Work experience placement
    Remote work

    Luma

    Redwood City, CA
    5 days ago
  •  ...Senior Site Reliability Engineer LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue. We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly... 
    Senior
    Full time
    Work at office
    Flexible hours
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  • $160k - $240k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe...  ...at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site Reliability...  ...Participate in on-call rotations and lead incident response activities; run and... 
    Senior

    Fiserv

    Sunnyvale, CA
    5 days ago
  • $132.6k - $214.5k

     ...you will collaborate closely with our engineering teams to develop innovative solutions that...  ...' performance and health. As a Senior Staff SRE with the Cortex Observability...  ...operability of the product and ensure the reliability and availability of our services. Qualifications... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  •  ...Lawyer / Attorney / Counsel - Junior / Senior / Lead Palo Alto, California, United States About the Job Catalyst Labs is a leading talent agency with a specialized vertical in Legal, Regulatory Compliance, and Corporate Governance. We operate deeply inside our... 
    Senior
    Contract work

    Catalyst Labs, LLC

    Palo Alto, CA
    4 days ago
  • $61k - $101k

     ...Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need...  ..., communities of practice, guilds, and conferences. We lead reuse-first adoption of AI-assisted reliability workflows across... 
    Senior
    Full time

    J.P. Morgan

    Palo Alto, CA
    9 days ago
  • Citi is seeking a Private Banker Sr. Principal in Palo Alto, California to effectively manage high-net-worth client relationships and serve as a trusted financial advisor. This role requires a broad understanding of financial strategies and the ability to identify client...
    Senior

    Citi

    Palo Alto, CA
    4 days ago
  • $226.78k - $293.48k

     ...Senior AI Storage Solutions Lead, Solutions Architecture Join us to do the best work of your career and make a profound social impact as a Senior...  ...assumptions and reimagining how work gets done. Engineers define intent, author precise specifications, and orchestrate... 
    Senior

    Socket.dev

    Santa Clara, CA
    3 days ago
  • $222k - $300.5k

     ...TeamIntuit's Infrastructure and Site Reliability organization owns the...  .... The Fintech Platform Systems Engineering team builds and operates the AWS...  ....The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  • $262k - $364k

     ...within the AViD ecosystem have reliability and uptime appropriate to...  ...and performance.Build creative engineering solutions to operations and infrastructure...  ...for production operations.Lead and contribute to the cross-...  ...in a strategic way.Site Reliability Engineering (SRE)... 
    Senior

    Google

    Mountain View, CA
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Work at office

    Chase

    Palo Alto, CA
    4 days ago
  •  ...Responsibilities & Expectations The Senior Team Leader is an experienced Executive Protection Agent tasked with leading a team of at least 5-15 Agents, wherein you will...  ...of the client as the trusted, senior most on-site leader. Scheduling, personnel management, critical... 
    Senior
    Local area
    Shift work
    Night shift

    Crisis24

    Palo Alto, CA
    2 days ago
  •  ...NVIDIA is seeking a Senior Software and System Architect to join the Networking Software Architecture group to lead architecture for cloud-networking and security solutions and to design state-of-the-art system architectures for DPUs and NICs. You will build end-to-end... 
    Senior
    Remote job

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $200k - $260k

     ...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical...  ...practical experience. ~8+ years of experience in a senior-level role within Site Reliability Engineering or similar... 
    Work at office
    Home office

    Glean - Mountain View, CA, US

    Mountain View, CA
    3 days ago
  • $217.57k - $260k

     ...explicitly states otherwise, all roles are on-site five days per week at one of our...  ...here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high...  ...experience operating at this scale and leading infrastructure through significant... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    ID.me

    Mountain View, CA
    4 days ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within... 

    Chase

    Palo Alto, CA
    5 days ago
  •  ...Consultant to deliver high-impact consulting and implementation services, guiding customers through complex HCM SaaS engagements. You will lead discovery, design, testing, and go-live readiness while partnering with cross-functional teams to ensure on-time, high-quality... 
    Senior

    Jobleads-US

    Mountain View, CA
    4 days ago
  •  ...AI-powered observability platform engineered for scale — ingesting logs, metrics...  ...ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer to own and drive...  ...posting on the Snowflake Careers Site for salary and benefits information... 
    Senior
    Full time

    Snowflake

    Menlo Park, CA
    1 day ago
  • $135k - $175k

     ...About the Role We are seeking an experienced and passionate Senior Software Engineer to join our dynamic engineering team. In this role, you will...  ...and APIs using Java and the Spring Boot framework. Lead & Mentor: Take ownership of major features from conception... 
    Senior
    Full time

    Lirvana Labs

    Menlo Park, CA
    1 day ago
  • $152k - $214k

     ...Design APIs and SDKs together with external teams Support engineering development lifecycle processes in a highly regulated and safety...  ...requirements and to test new ideas Mentor other engineers Lead technical discussions, feature development, and architecture reviews... 
    Senior
    Full time

    Drivemode

    Mountain View, CA
    1 day ago
  • $145k - $182k

     ...the world to deliver code to their users reliably, efficiently, securely and quickly,...  ..., Security Testing Orchestration, Chaos Engineering, Software Engineering Insights and continues...  ...specifications, designs, and code Work alongside Site Reliability Engineers and cross... 
    Senior
    Full time
    Local area
    Immediate start
    Flexible hours

    Harness

    Mountain View, CA
    1 day ago
  • $166k - $244k

     ...we’re a team of scientists, engineers, machine learning experts and...  ...set you up for success as an Senior Software Engineer in the Gemini...  ...their user experience  Lead the full-stack implementation...  ...Strong communication skills and a reliable team player Proven... 
    Senior
    Full time

    Deepmind

    Mountain View, CA
    1 day ago
  • $220k - $230k

     ...market leaders for our proven ability to generate value and unlock opportunities that were previously unattainable.  The Senior Software Engineer (CALC engine) is a specialised role working on a complex Java based high-performance OLAP engine. Our engine is a multi-... 
    Senior
    Full time
    Work at office
    Flexible hours

    Aera Technology

    Mountain View, CA
    1 day ago
  •  ...running. Location: 5 on-site days a week in Sunnyvale, CA...  ...Our Team's Vision: Our Engineering team is shaping the future of...  ...looking for an experienced Senior Site Reliability Engineer (SRE) with a strong...  ...and infrastructure updates Lead incident response and... 
    Senior
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    1 day ago
  • $307.32k

     ...started. Our scientists, engineers, sales executives, and...  ...a high-performing Senior Software Engineer to...  ...optimize performance and reliability, enhancing overall...  ...facilities) Daily on-site lunches provided from...  ...on top of (4) industry leading company benefits (free... 
    Senior
    Full time
    Worldwide
    2 days per week
    3 days per week

    Billiontoone

    Menlo Park, CA
    1 day ago
  •  ...Senior Technical Marketing Engineer NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It...  ...are looking for a Senior Technical Marketing Engineer to lead and develop technical content for DSX. Want to create and... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv. Sr. Site Reliability Engineer About Clover Clover is a pioneer...  ...confidence. What Does A Successful Senior Site Reliability Engineer Do At Fiserv... 
    Senior
    Worldwide

    BentoBox

    Sunnyvale, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Lead Site Reliability Engineer. Be the first to apply!