Staff Site Reliability Engineer (Production Engineer)
$119k - $170kZscaler
Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will be responsible for all aspects of the Zscaler production data center services, including servers, operating systems, storage, and supporting systems. You will be an instrumental part of the Site Reliability Engineering team, ensuring the availability, latency, performance, efficiency, and scalability of a cloud that processes tens of billions of transactions daily. What you’ll do (Role Expectations) Own the reliability of a large‑scale cloud service (Linux/BSD, bare metal, Kubernetes, custom load balancing, SD‑WAN) by partnering with Engineering and Network teams to define requirements early, conduct operability reviews, and contribute code/design docs for platform resilience Develop and operate end‑to‑end observability (metrics/logs/traces, dashboards, alerting) and incident tooling to manage SLOs/error budgets, reduce noise, and improve system detection and diagnosis Participate in an on‑call rotation to lead full‑cycle incident response; perform deep cross‑stack troubleshooting (OS, networking, distributed systems, packet captures, core dumps) to drive permanent software fixes and codify learnings into runbooks and tests Build and maintain everything‑as‑code for fleet and service lifecycle, driving provisioning, configuration, release automation, canary deployments, and complex rollout/rollback workflows Continuously improve platform hygiene through consistent OS/app upgrades, dependency/vulnerability patching, capacity and performance tuning, and strict CI/CD validation prior to production rollouts Who You Are (Success Profile) You thrive in ambiguity. You’re comfortable building the path as you walk it. You thrive in a dynamic environment, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high‑level strategy and hands‑on execution. You are a problem‑solver. You love running towards the challenges because you are laser‑focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high‑trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback—knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We’re Looking for (Minimum Qualifications) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI‑driven solutions to optimize outcomes within your functional domain US Citizenship is required (due to the nature of assigned customers) 5+ years industry experience in software engineering, infrastructure software, and/or platform engineering Proficiency in at least one programming language (such as Python, Bash, or Go) with demonstrated ability to write production‑quality code (testing, code reviews, CI, maintainable design, scripting for diagnostics) Strong Linux/Unix systems fundamentals (process/memory, filesystems, networking stack basics, debugging/perf troubleshooting) and solid understanding of networking protocols and components (e.g., DNS, TCP/IP, ICMP, OSI model, subnetting, and load balancing/traffic concepts) Proven experience operating production services (including incident response, troubleshooting, reducing toil) and managing BSD in production to drive systemic fixes through platform engineering What Will Make You Stand Out (Preferred Qualifications) Experience leveraging AI/ML frameworks or AIOps tools to build predictive anomaly detection, automate root‑cause analysis, and optimize large‑scale infrastructure reliability Proven expertise in operating Kubernetes at scale Deep experience with the Prometheus/OpenTelemetry ecosystems, including instrumenting golden signals, defining SLOs, and performing alert tuning to ensure high‑availability environments Benefits Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In‑office perks, and more! Salary Base Pay Range: $119,000—$170,000 USD Equal Employment Opportunity Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy‑related support. #J-18808-Ljbffr Zscaler
$127k - $249k
MongoDB is seeking a Site Reliability Engineer (SRE) with over 5 years of experience to join their Atlas team in Boston. This hybrid role involves maintaining and growing the Atlas platform, developing a multi-cloud environment for business-critical applications, and collaborating...SuggestedFlexible hours$75.7k - $136.3k
...automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and...SuggestedWork experience placementWork at office$140k - $210.9k
...the marketplace for new products and services more... ...opportunities for FRFS staff. The Federal Reserve... ...will be primarily on-site with residency commutable... ...or software engineering backgrounds (e.g., Java... ...operating and improving reliability of distributed production...SuggestedFull timeTemporary workPart timeWork at officeShift work- ...Job Title: Site Reliability Engineer Location: Remote with Quarterly visits to Chennai, Tamil Nadu, India Duration: Full-Time bout BigRio: BigRio is a remote-based, technology consulting firm headquartered in Boston, MA. We deliver software solutions...SuggestedFull timeRemote work
$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...rollback procedures Collaborating with product engineering teams to troubleshoot...SuggestedPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$104.9k - $174.7k
...distributed and fault-tolerant systems within agreed reliability objectives, whilst enabling the fast flow of feature and... ...skills. About team; This diverse team of Engineers in assisting multiple product teams as we continue to innovate all of our products within...Local areaImmediate startWorldwide$146.4k - $263.6k
...enjoy working with a diverse multi-national team of engineering talents? Join our highly skilled Site Reliability team Our team designs, develops, and... ...and infrastructure that support Akamai's Compute products and services. We specialize in building and maintaining...Work experience placementWork at office$121.4k - $218.6k
...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...spanning the globe. You'll collaborate with product teams from the earliest stages of... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for:...Work experience placementWork at office$51.9 per hour
...job is responsible for the reliability, availability, and performance... ...This role blends software engineering, clinical engineering, and... ...-functionally with AHN site leaders and teams to navigate... ...performance management and staff productivity.Plan, organize, staff, direct...For contractorsLocal area$51.9 per hour
...Company: Allegheny Health Network Job Title: Site Reliability Engineering – Clinical & Facility Services General Overview This role ensures the... ...Environment of Care (EOC), supporting patients, providers, and staff. The position blends software engineering, clinical...Local area- ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform... ...efforts. The results of your work will empower our product engineers to move more autonomously in the build of our features...Shift work
$121.5k - $306.4k
...and provides input on best practices for reliability and functionality. Establishes direction... ..., executing improvements, building site reliability knowledge, and providing clear... ..., as well as reflect Oracle's differing products, industries and lines of business. Candidates...Temporary workFlexible hours$84.9k - $209.5k
...Intelligence Platform. This team will focus on product development and product strategy for... ...your contribution to make it a special engineering center with the focus on excellence.... ...yet been documented as SOPs for Level1 staff. You will usually get called in during major...Temporary workImmediate startFlexible hours$160k - $225k
...sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification... .... About the Role Manifold is looking for a Staff Site Reliability Engineer (SRE) to work at the intersection of AI, data infrastructure...$75.2k - $95.3k
...looking for a highly motivated and high‑potential entry‑level Site Reliability Engineer (SRE) to join our team and help drive meaningful business... ...part of the SRE transformation at WEX. Our sophisticated products power a diverse range of customer businesses, and the...Work experience placementFlexible hours- ...Education Desired: Bachelor of Computer Engineering Travel Percentage: 0% We are FIS. Our... ...champion diversity to deliver the best products and solutions for our colleagues,... .../desire to improve application systems reliability and automate manual support tasks, to facilitate...Full timeWork at officeRemote workWork from homeFlexible hours
- A leading fintech company is seeking an experienced technical support specialist in Boston, MA. You will troubleshoot production system issues and provide automation support for financial applications. The role requires at least 5 years of experience in Java and database...Remote jobWork at office
$130k - $150k
Site Reliability Engineer - Disaster Recovery & Business Continuity Boston, MA, United States; Chicago... ...Security Information Technology staff are based in the Boston, Chicago, London... ...culture, and shared ownership of production outcomes. Experience operating and improving...Work at officeWork from home3 days per week- ...meet the needs of the marketplace for new products and services more quickly, seek to... ...our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team,... ...someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and...Full time
$150k - $185k
...error-proofing processes and boosting productivity, capturing and analyzing real-time data... ...work best. You enjoy building for other engineers equally, if not more, than building for... ...best practices, SLIs/SLOs, and reliability culture across engineering teams. Help...Temporary workWork at officeFlexible hours3 days per week- Site Reliability Engineer at the organization. Key technologies: Kubernetes, Prometheus, Grafana. Key Responsibilities Define and track SLOs, SLIs and error budgets Design and implement observability stacks (metrics, logging, tracing) Automate toil and improve system...
$95k - $171k
A leading cloud computing company seeks a Site Reliability Engineer II to join their Inference Cloud Team. The role involves building dashboards, writing automation in Python or Go, and collaborating with engineering teams to ensure AI infrastructure reliability. Candidates...Flexible hours$75.2k - $95.3k
WEX, Inc. is seeking a highly motivated entry-level Site Reliability Engineer (SRE) in Boston, MA. This role is integral to our SRE transformation, contributing to the performance and reliability of our systems. As an SRE, you'll monitor system reliability, manage incidents...$127k - $249k
The Team Platform Engineering is the department within SRE that is responsible for a range... ...our developers to build and ship products to delight our customers. We manage the... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...Work at officeLocal areaRemote workWorldwideFlexible hours$128k - $160k
...driver of progress, come build the future together. The Crown Is Yours As a Senior Site Reliability Engineer, you'll build and scale the critical infrastructure behind every product. In this role, you'll take on complex challenges across global data centers, multiple...Full timeImmediate start$146.4k - $263.6k
...large datasets to analyze and measure the performance and reliability of our platform. We are networking data scientists: we... ...project‑specific work. Driving partnership with Engineering, Operations and Product teams to help guide adoption of new features and processes...Work at office$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and... ...our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is...Local areaWorldwideFlexible hours$150k
...daily at Coalition. About The Role We are looking for a Staff Site Reliability Engineer to lead AI enablement across our engineering... ...to ensure AI‑generated output is reliable, secure, and production‑worthy. This role owns that layer. This role blends building...Fixed term contractWork experience placementWork at officeRemote workHome officeFlexible hoursShift work$127k - $249k
...are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain... ...Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure... ...everything we do results in a stronger product and a better experience for all Atlas...Local areaRemote workFlexible hours$140k - $210.9k
The Federal Reserve Bank of Boston seeks a Senior Site Reliability Engineer to enhance their payments service, FedNow. This position involves significant responsibilities in operating the production environment and the design of automation and monitoring tools to ensure...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer (Production Engineer). Be the first to apply!
- assistant chief engineer Boston, MA
- technology administrator Boston, MA
- project engineer assistant project manager Boston, MA
- engineering aide Boston, MA
- senior staff systems engineer Boston, MA
- assistant electrical engineer Boston, MA
- staff design engineer Boston, MA
- assistant engineering manager Boston, MA
- software engineer staff Boston, MA
- senior staff engineer Boston, MA

