Senior Site Reliability Engineer
Bank of America ATM
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve. Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations. At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!This job is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing observability designs through instrumentation and dashboards, identifying root causes of complex/impactful issues, partnering with cross functional teams to deliver sustainable design patterns, and driving early adoption of non-functional production support requirements. Job expectations include automating services to improve reliability and efficiency and influencing a culture of innovation and continuous improvement.
Position Summary:
The Senior GCP Site Reliability Engineer acts as an advanced senior individual contributor responsible for designing, implementing, and maturing reliability engineering capabilities. The role focuses on complex technical problem solving, reliability architecture, automation strategy, observability maturity, platform resiliency, and operational excellence. The ideal candidate can operate across both strategy and execution: defining patterns, building automation, mentoring engineers, resolving hard platform issues, and driving long-term reliability improvements.
Key Responsibilities:
- Serve as a senior technical authority for GCP platform reliability, resiliency, observability, automation, and production readiness
- Design and implement advanced reliability patterns for Azure landing zones, private networking, DNS, firewalls, regional resiliency, service health checks, and workload onboarding
- Lead complex platform reliability initiatives such as secondary-region readiness, egress/ingress observability, private DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation
- Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting for Azure platform services
- Develop reusable Terraform modules, automation frameworks, and CI/CD patterns that improve consistency, compliance, and operational quality
- Drive observability improvements using Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools
- Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, automation opportunities, and operational controls
- Lead deep technical investigations for major incidents, recurring problems, platform defects, or service degradation events
- Partner with security and governance teams to integrate IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness into Azure platform operations
- Provide technical design input for new Azure services and workloads to ensure operational readiness before production adoption
- Mentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production support
- Create executive-ready technical summaries, reliability narratives, and recommendations for leadership review
- Establish reusable standards for runbooks, dashboards, health checks, problem records, post-incident reviews, and production-readiness gates
- Designs solutions to visualize key production support metrics enabling Operational Readiness and Site Reliability Engineer teams to identify scenarios requiring intervention
- Develops software solutions and/or improved processes to address work identified as ‘toil’ by collaborating with key partners to identify, track and remediate processes to free time allocated to reliability
- Partners with Development and Infrastructure teams to create error budget policies prioritizing reliability stories that fall below Service Level Objective (SLO) thresholds and suggests code optimizations, additional instrumentation and/or logging structures to gain service reliability visibility
- Identifies and plans for capacity bottlenecks, vulnerabilities and opportunities for reliability improvement, such as low level error rates and 'noise', and reduces manual support effort and/or improves system reliability
- Assesses monitoring for new changes with development partners and works with monitoring tools team to monitor dashboards and enhance application and system monitoring designs
- Engages as a subject matter expert in incident triage efforts, failure scenario modelling and works with the Problem Manager to diagnose root causes for complex/high impact incident/problem management investigations
- Collaborates with Development and Infrastructure teams to understand technical solutions and develop Service Level Indicators and SLOs to measure/improve the reliability of the services they support
- see position summary required/desired qualifications
Required Qualifications:
- 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
- Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise for infrastructure provisioning
- Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code
- Experience developing and maintaining Terraform modules and infrastructure configurations to support automated cloud environments
- Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes
- Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows
- Good understanding of GCP services and cloud architecture fundamentals, including networking (VPCs, IAM, load balancing)
- Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
- Experience supporting automation and standardization efforts to improve consistency and efficiency in cloud deployments
- Understanding of monitoring, logging, and observability tools to support system reliability and performance
- Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
Desired Qualifications:
- Ability to collaborate effectively with engineering, architecture, and security teams to support reliable and secure platform operations
- Strong problem-solving and analytical skills, with the ability to diagnose and resolve infrastructure issues
- Effective communication skills, with the ability to work within cross-functional teams and document technical solutions clearly
- Interest in learning and applying emerging technologies and automation techniques (including AI/ML where applicable) to improve platform reliability
Skills:
- Architecture
- Collaboration
- Innovative Thinking
- Result Orientation
- Solution Design
- Adaptability
- Analytical Thinking
- Influence
- Stakeholder Management
- Technical Strategy Development
Shift:
1st shift (United States of America)Hours Per Week:
40- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...SeniorTemporary workLocal area
- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Chief Data & Analytics...SeniorWork at office
$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online...SeniorFull timeTemporary workLocal areaRemote workWeekend work- ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security...Senior
- LiveRamp, a data collaboration platform leader, is seeking a Senior Site Reliability Engineer in San Francisco with 5+ years of SRE/DevOps experience. The role focuses on deployments, 24/7 support across regions, and establishing SRE best practices. Strong skills in Terraform...Senior
$174k - $252k
Senior Software Engineer, Site Reliability Engineering corporate_fare Google place Seattle, WA, USA ; Kirkland, WA, USA Mid Experience driving progress, solving problems, and mentoring more junior team members; deeper expertise and applied knowledge within relevant area...SeniorTemporary work- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex...Work at office
- ...The selected colleague will work at an MUFG office or client sites four days per week and work remotely one day. A member of... ...MUFG is seeking a highly motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable, and reliable web application environment...Full timeWork at officeLocal areaRemote work
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer focussed on Operations Excellence, you will have the opportunity to shape how we respond to, learn from, and...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Cloud Foundational Services team, you hold a leadership role in your team, demonstrate...
- ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new opportunities...
- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Consumer and Investment Banking team, you will solve complex and broad business problems...Local area
- ...QualificationsBA degree or higherComputer Science or related disciplines.7+ years working with technical teams to define software business requirementsSummaryFunction: Information TechnologyExperience level: Mid-Senior LevelIndustry: Information Technology And ServicesSenior
$137.75k - $185k
...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Asset and Wealth Management team, you will solve complex and broad business problems...Full timeWork experience placement- ...we serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted... ...scalability, and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across...Remote workFlexible hours
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will... ...Professional development From entry-level employees to senior leaders, we believe there’s always room to learn. We offer...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...Job Description Job Description BCforward is currently seeking a highly motivated SRE Software Engineer. Job Title: SRE Software Engineer Location: Jersey City, NJ Duration: Temp - 12 months Job Description We are seeking a Software Engineer-Other...Temporary work
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Temporary work
- Job ID: 25677433Reference Number: 25-00616Title: Senior Python DeveloperLocation: jersey city, NJJob Type: Full Time/ContractPosted... ...developers with 6+ years' experience Experience with PySpark Should also have a strong SQL background and have data engineering experienceSeniorFull time
- Role SummaryWe are seeking a Senior Java MuleSoft Engineer P4 with strong expertise in enterprise integration API development and cloudbased integration platforms The ideal candidate will have handson experience in MuleSoft Anypoint Platform Java Spring Boot APIled architecture...Senior
$157k - $235k
...Talkdesk is looking for a Senior Solutions Engineer specializing in Banking, Financial Services, and Insurance (FSI). You will engage throughout the sales lifecycle, offering industry-centered expertise, conducting research, and designing impactful presentations. The role...Senior- ...collaborative company where innovation, creativity, ownership, and impact are part of everyday work. ABOUT THE ROLE:As a Senior Software Engineer at Global-e, you will design and deliver the core services behind our global logistics platform. You will drive innovation...SeniorWork at officeWorldwide
- ...focused on improving the security, release reliability, maintainability, and operational... ...critical banking applications.Forward Deployed Engineers will conduct targeted, hands-on... ...rapid release cadence.The role combines senior software engineering, DevOps, SRE, test...Senior
$150k - $180k
Senior Software EngineerLocation: Remote - US | Team: Engineering - Enterprise Learning & SkillsAbout the role We're looking for a Senior Full-Stack Software Engineer to join the Enterprise Learning & Skills (ELS) engineering team and help build, scale, and evolve a large...SeniorFull timeRemote work- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the AI Machine Learning and Data platform team, you hold a leadership role in...
$108k - $216k
...PermanentCompany: WalmartBusiness Segment: Home OfficeRole summary: The Senior Software Engineer will lead the delivery of scoped features and models... ..., uphold engineering excellence, and support system reliability while mentoring peers and contributing to a high-...SeniorFull timeTemporary workPart time- ...to understand business requirements and translate them into technical requirements ~ Demonstrate a solid understanding of core engineering principles ~ Familiarity with modern software engineering practices and continuous integration and delivery ~ Comfortable working...SeniorRemote jobFull timeHome office
$140k - $200k
...asked to participate in an on-site interview as part of the... ...every collector, old or new. Our engineering mission is to democratize technology... ....We’re looking for a Senior Platform Engineer to join our... ...observability), improving the reliability, scalability, and cost-efficiency...SeniorFull timeRemote workWorldwideFlexible hoursShift work- ...leader in catastrophe-exposed property insurance, is seeking a Senior Software Engineer to join our Policy Onboarding and Home Services team. As a... ...after they purchase a policy, ensuring a seamless, reliable, and modern experience. You'll contribute to architectural...SeniorFor contractorsLive inRemote work
- Job ID: 29181067Reference Number: 26-00872Title: Senior Java Engineer | Jersey City, NJ (Onsite)Posted Date: 2026-08-28Contact: Alok Kumar VermaContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Local to New Jersey or Neighbour States...SeniorLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Jersey City, NJ
- site reliability engineer sre Jersey City, NJ
- senior network engineer remote Jersey City, NJ
- senior app developer Jersey City, NJ
- senior manager legal Jersey City, NJ
- sr project manager Jersey City, NJ
- senior commercial counsel Jersey City, NJ
- senior data manager Jersey City, NJ
- senior manager tax Jersey City, NJ
- senior construction accountant Jersey City, NJ



