Site Reliability Engineer
J.P. Morgan
hackajob is collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms and Foundational Services (IPFS) team, you work with your fellow stakeholders to define non-functional requirements (NFRs) and availability targets for the services in your application and product lines. You will ensure those NFRs are accounted for in your products' design and test phases, that your service level indicators are effectively measuring customer experience, and that service level objectives are defined with stakeholders and implemented in production. Job Responsibilities Creates and delivers high quality designs, roadmaps, and program charters alongside the engineering team Acts as a key resource and mentor for technologists in your area seeking advice on technical and business issues, and serves as a culture carrier and site reliability adoption champion for your team Collaborates with others to create and implement observability and reliability designs for complex systems which are robust, stable, and do not incur additional toil or technical debt Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements. Drives evolution and debugging of critical components by understanding application and platform interdependencies and limitations Provides comprehensive and ongoing guidance, tools, and solutions to support the firms' growth Make significant contributions to JPMorganChase's site reliability community via internal forums, communities of practice, guilds, and conferences Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls. Required qualifications, capabilities, and skills Formal training or certification on site reliability engineering concepts and 5 years applied experience Advanced knowledge in site reliability culture and principles with demonstrated ability to implement site reliability within an application or platform Advanced knowledge and experience in observability such as white and black box monitoring, service level objectives, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Expert-level proficiency in Java, Go (Golang), Python, and Terraform for building enterprise-grade applications, high-performance systems, automation, and infrastructure as code Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity. Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes. Advanced knowledge of software applications and technical processes with considerable depth in multiple technical disciplines including distributed systems, microservices architecture, and cloud-native technologies Hands-on experience building AI Agents and autonomous systems with proficiency in AI frameworks (LangChain, LangGraph, AutoGen, CrewAI) and leveraging AI development tools (GitHub Copilot, Claude, etc.) to accelerate development and innovation and Expertise in designing and implementing logging pipelines (Fluentd, Logstash, Vector) and systems for metrics collection, analysis, and distributed tracing Strong experience building production-grade RESTful APIs and designing message queue architectures (Kafka, RabbitMQ, SQS) for event-driven systems; and expertise in graph databases (Neo4j, TigerGraph), vector databases (Pinecone, Weaviate, Chroma), and integrating multiple data stores for AI-powered systems Proficiency with containerization (Docker, Kubernetes), CI/CD pipelines, and GitOps workflows Ability to communicate data-based solutions with complex reporting and visualization methods, recognized as an active contributor of the engineering community, and continues to expand network and leads evaluation sessions with vendors to see how offerings can fit into the firm's strategy Preferred qualifications, capabilities, and skills Experience with MCP (Model Context Protocol) Servers or similar agent frameworks for building autonomous systems, and understanding of LLM integration, prompt engineering, and RAG (Retrieval-Augmented Generation) Familiarity with AI/ML model building, deployment, and lifecycle management using frameworks like TensorFlow, PyTorch, or scikit-learn Experience with big data technologies (Hadoop, Spark, Flink), analytical databases, NoSQL databases (MongoDB, Cassandra, DynamoDB), and time-series databases (InfluxDB, TimescaleDB) Knowledge of security best practices and compliance requirements in highly regulated industries, with experience in chaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos) and GameDay exercises Contributions to open-source projects, particularly in SRE, observability, or AI/ML domains, and certifications in cloud platforms (AWS, Azure, GCP) Strong communication skills with ability to mentor and educate others on site reliability principles and practices, and ability to anticipate, identify, and troubleshoot defects found during testing This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase's review of criminal conviction history, including pretrial diversions or program entries. ABOUT US JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world's most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management. We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation. JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans ABOUT THE TEAM Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we're setting our businesses, clients, customers and employees up for success.aa415a4b-8b21-40fc-a65c-70d2b25ca29a
$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...SuggestedFull time- ..., Elise AI, IBM and Accern. Position Summary We are hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep technical execution. You will be responsible for defining reliability...SuggestedShift work
- ...ActiveHours is looking for an experienced DevOps Engineer to enhance platform automation and site reliability in a collaborative environment. The role involves automating key systems, improving visibility through metrics, and troubleshooting critical problems. You'll...Suggested
- ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments...Suggested
- ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data...Suggested
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Work at office
$90k - $180k
...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division. We...Full timeRemote workShift work$217.57k - $260k
...job description explicitly states otherwise, all roles are on-site five days per week at one of our offices in McLean, VA;... ...which can be found here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale infrastructure...Full timeTemporary workWork at officeRemote workFlexible hoursShift work- ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within...
$145k - $165k
...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to...Work at officeImmediate start$145k - $165k
...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...- ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...
- ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable...
- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud...Full time
- ...CoreWeave seeks a Reliability Lead for Common Services in Sunnyvale, CA. You will establish and lead reliability engineering, production operations, and an observability-driven culture across multiple teams. Partner with engineering leaders to define SLOs/SLIs, drive...
$135.6k - $180k
...operations team. This role involves overseeing 24/7 operational stability, enhancing processes and systems, and mentoring a diverse engineering team. The ideal candidate will have over 8 years of technical operations experience, proficiency in infrastructure automation...$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of...Full timeContract workWorldwide$230k - $250k
...Site Reliability Engineer Forward is transforming how the world's most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry's first network digital twin — a mathematically precise model of the production network that...Night shift- ...Site Reliability Engineer (SRE) The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must be a U.S. Person (i.e. U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee). This position may also perform...Permanent employmentWorldwideShift work
$264.52k
...office, 2 days remote - Must be able to report to local office. REQUIREMENTS: Master’s degree in Computer Science, Mechatronics Engineering, Software Engineering, Electrical Engineering, Computer Engineering, or related field of study. Five (5) years of experience as a...Full timeWork at officeLocal areaRemote workWork from home$165k - $190k
...DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-... ...security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with...Work from homeFlexible hours- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...
- ...Job Description Job Description Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes...
$85k - $90k
...things. And we will be 100% committed to helping you reach your full potential. Job Description General Description: The Reliability Test Engineer position is responsible for executing reliability testing, operating lab equipment and thermal chambers, documenting...Full timeWork experience placementLocal areaImmediate startRelocation- ..., and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...Remote work
$148k - $222k
...accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a...Permanent employmentFull timeTemporary work$185k - $278k
...transforming them into scalable solutions. * Debugging OS and engineering issues within our provided Linux environment. *... ...efficiently. The Impact You Will Have: * Enhancing the reliability and performance of our engineering environment. * Streamlining...Remote work$180k - $260k
...effortless integration into customers' logistics operations. About the role We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you will work...Odd jobWork at officeRemote work- ...Kubernetes operators. We are building out several platform engineering teams that together own the full stack, from managed bare metal... ...resources. Storage: provisioning workflows, attachment reliability, performance tuning, and failure handling. Confidential computing...Full timeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- construction site safety Palo Alto, CA
- on-site clinical research associate (traveling/remote) Palo Alto, CA
- site safety Palo Alto, CA
- historic site Palo Alto, CA
- IT site lead Palo Alto, CA
- site leader Palo Alto, CA
- junior website developer Palo Alto, CA
- official site Palo Alto, CA

