Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer (Production Engineer)

$119k - $170k

Zscaler

Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will be responsible for all aspects of the Zscaler production data center services, including servers, operating systems, storage, and supporting systems. You will be an instrumental part of the Site Reliability Engineering team, ensuring the availability, latency, performance, efficiency, and scalability of a cloud that processes tens of billions of transactions daily. What you’ll do (Role Expectations) Own the reliability of a large‑scale cloud service (Linux/BSD, bare metal, Kubernetes, custom load balancing, SD‑WAN) by partnering with Engineering and Network teams to define requirements early, conduct operability reviews, and contribute code/design docs for platform resilience Develop and operate end‑to‑end observability (metrics/logs/traces, dashboards, alerting) and incident tooling to manage SLOs/error budgets, reduce noise, and improve system detection and diagnosis Participate in an on‑call rotation to lead full‑cycle incident response; perform deep cross‑stack troubleshooting (OS, networking, distributed systems, packet captures, core dumps) to drive permanent software fixes and codify learnings into runbooks and tests Build and maintain everything‑as‑code for fleet and service lifecycle, driving provisioning, configuration, release automation, canary deployments, and complex rollout/rollback workflows Continuously improve platform hygiene through consistent OS/app upgrades, dependency/vulnerability patching, capacity and performance tuning, and strict CI/CD validation prior to production rollouts Who You Are (Success Profile) You thrive in ambiguity. You’re comfortable building the path as you walk it. You thrive in a dynamic environment, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful. You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome. True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high‑level strategy and hands‑on execution. You are a problem‑solver. You love running towards the challenges because you are laser‑focused on finding the solution, knowing that solving the hard problems delivers the biggest impact. You are a high‑trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback—knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust. You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose. What We’re Looking for (Minimum Qualifications) Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI‑driven solutions to optimize outcomes within your functional domain US Citizenship is required (due to the nature of assigned customers) 5+ years industry experience in software engineering, infrastructure software, and/or platform engineering Proficiency in at least one programming language (such as Python, Bash, or Go) with demonstrated ability to write production‑quality code (testing, code reviews, CI, maintainable design, scripting for diagnostics) Strong Linux/Unix systems fundamentals (process/memory, filesystems, networking stack basics, debugging/perf troubleshooting) and solid understanding of networking protocols and components (e.g., DNS, TCP/IP, ICMP, OSI model, subnetting, and load balancing/traffic concepts) Proven experience operating production services (including incident response, troubleshooting, reducing toil) and managing BSD in production to drive systemic fixes through platform engineering What Will Make You Stand Out (Preferred Qualifications) Experience leveraging AI/ML frameworks or AIOps tools to build predictive anomaly detection, automate root‑cause analysis, and optimize large‑scale infrastructure reliability Proven expertise in operating Kubernetes at scale Deep experience with the Prometheus/OpenTelemetry ecosystems, including instrumenting golden signals, defining SLOs, and performing alert tuning to ensure high‑availability environments Benefits Various health plans Time off plans for vacation and sick time Parental leave options Retirement options Education reimbursement In‑office perks, and more! Salary Base Pay Range: $119,000—$170,000 USD Equal Employment Opportunity Zscaler is committed to providing equal employment opportunities to all individuals. We strive to create a workplace where employees are treated with respect and have the chance to succeed. All qualified applicants will be considered for employment without regard to race, color, religion, sex (including pregnancy or related medical conditions), age, national origin, sexual orientation, gender identity or expression, genetic information, disability status, protected veteran status, or any other characteristic protected by federal, state, or local laws. See more information by clicking on the Know Your Rights: Workplace Discrimination is Illegal link. Zscaler complies with all applicable federal, state, and local pay transparency rules. Zscaler is committed to providing reasonable support (called accommodations or adjustments) in our recruiting processes for candidates who are differently abled, have long term conditions, mental health conditions or sincerely held religious beliefs, or who are neurodivergent or require pregnancy‑related support. #J-18808-Ljbffr Zscaler

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer (Production Engineer) in Boston, MA vacancy
  • $127k - $249k

    MongoDB is seeking a Site Reliability Engineer (SRE) with over 5 years of experience to join their Atlas team in Boston. This hybrid role involves maintaining and growing the Atlas platform, developing a multi-cloud environment for business-critical applications, and collaborating... 
    Suggested
    Flexible hours

    MongoDB

    Boston, MA
    2 days ago
  • $75.7k - $136.3k

     ...automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Boston, MA
    3 days ago
  • $140k - $210.9k

     ...the marketplace for new products and services more...  ...opportunities for FRFS staff. The Federal Reserve...  ...will be primarily on-site with residency commutable...  ...or software engineering backgrounds (e.g., Java...  ...operating and improving reliability of distributed production... 
    Suggested
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve System

    Boston, MA
    1 day ago
  •  ...Job Title: Site Reliability Engineer Location: Remote with Quarterly visits to Chennai, Tamil Nadu, India Duration: Full-Time bout BigRio: BigRio is a remote-based, technology consulting firm headquartered in Boston, MA. We deliver software solutions... 
    Suggested
    Full time
    Remote work

    Saviance

    Boston, MA
    1 day ago
  • $95k - $171k

     ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:...  ...rollback procedures Collaborating with product engineering teams to troubleshoot... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Boston, MA
    2 days ago
  • $104.9k - $174.7k

     ...distributed and fault-tolerant systems within agreed reliability objectives, whilst enabling the fast flow of feature and...  ...skills. About team; This diverse team of Engineers in assisting multiple product teams as we continue to innovate all of our products within... 
    Local area
    Immediate start
    Worldwide

    RELX

    Cambridge, MA
    3 days ago
  • $146.4k - $263.6k

     ...enjoy working with a diverse multi-national team of engineering talents? Join our highly skilled Site Reliability team Our team designs, develops, and...  ...and infrastructure that support Akamai's Compute products and services. We specialize in building and maintaining... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    4 days ago
  • $121.4k - $218.6k

     ...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure...  ...spanning the globe. You'll collaborate with product teams from the earliest stages of...  ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for:... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    1 day ago
  • $51.9 per hour

     ...job is responsible for the reliability, availability, and performance...  ...This role blends software engineering, clinical engineering, and...  ...-functionally with AHN site leaders and teams to navigate...  ...performance management and staff productivity.Plan, organize, staff, direct... 
    For contractors
    Local area

    Highmark Health

    Boston, MA
    4 days ago
  • $51.9 per hour

     ...Company: Allegheny Health Network Job Title: Site Reliability Engineering – Clinical & Facility Services General Overview This role ensures the...  ...Environment of Care (EOC), supporting patients, providers, and staff. The position blends software engineering, clinical... 
    Local area

    Highmark Health

    Boston, MA
    3 days ago
  •  ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform...  ...efforts. The results of your work will empower our product engineers to move more autonomously in the build of our features... 
    Shift work

    Roberts Recruiting

    Boston, MA
    4 days ago
  • $121.5k - $306.4k

     ...and provides input on best practices for reliability and functionality. Establishes direction...  ..., executing improvements, building site reliability knowledge, and providing clear...  ..., as well as reflect Oracle's differing products, industries and lines of business. Candidates... 
    Temporary work
    Flexible hours

    Oracle

    Boston, MA
    3 days ago
  • $84.9k - $209.5k

     ...Intelligence Platform. This team will focus on product development and product strategy for...  ...your contribution to make it a special engineering center with the focus on excellence....  ...yet been documented as SOPs for Level1 staff. You will usually get called in during major... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Boston, MA
    4 days ago
  • $160k - $225k

     ...sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification...  .... About the Role Manifold is looking for a Staff Site Reliability Engineer (SRE) to work at the intersection of AI, data infrastructure... 

    Manifold AI

    Cambridge, MA
    1 day ago
  • $75.2k - $95.3k

     ...looking for a highly motivated and high‑potential entry‑level Site Reliability Engineer (SRE) to join our team and help drive meaningful business...  ...part of the SRE transformation at WEX. Our sophisticated products power a diverse range of customer businesses, and the... 
    Work experience placement
    Flexible hours

    WEX, Inc.

    Boston, MA
    5 days ago
  •  ...Education Desired: Bachelor of Computer Engineering Travel Percentage: 0% We are FIS. Our...  ...champion diversity to deliver the best products and solutions for our colleagues,...  .../desire to improve application systems reliability and automate manual support tasks, to facilitate... 
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours

    Dev

    Boston, MA
    5 days ago
  • A leading fintech company is seeking an experienced technical support specialist in Boston, MA. You will troubleshoot production system issues and provide automation support for financial applications. The role requires at least 5 years of experience in Java and database... 
    Remote job
    Work at office

    Dev

    Boston, MA
    5 days ago
  • $130k - $150k

    Site Reliability Engineer - Disaster Recovery & Business Continuity Boston, MA, United States; Chicago...  ...Security Information Technology staff are based in the Boston, Chicago, London...  ...culture, and shared ownership of production outcomes. Experience operating and improving... 
    Work at office
    Work from home
    3 days per week

    Charles River Associates

    Boston, MA
    5 days ago
  •  ...meet the needs of the marketplace for new products and services more quickly, seek to...  ...our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team,...  ...someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and... 
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    2 days ago
  • $150k - $185k

     ...error-proofing processes and boosting productivity, capturing and analyzing real-time data...  ...work best. You enjoy building for other engineers equally, if not more, than building for...  ...best practices, SLIs/SLOs, and reliability culture across engineering teams. Help... 
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    Tulip Interfaces

    Somerville, MA
    2 days ago
  • Site Reliability Engineer at the organization. Key technologies: Kubernetes, Prometheus, Grafana. Key Responsibilities Define and track SLOs, SLIs and error budgets Design and implement observability stacks (metrics, logging, tracing) Automate toil and improve system... 

    Gravity Engineering Services Pvt Ltd.

    Boston, MA
    5 days ago
  • $95k - $171k

    A leading cloud computing company seeks a Site Reliability Engineer II to join their Inference Cloud Team. The role involves building dashboards, writing automation in Python or Go, and collaborating with engineering teams to ensure AI infrastructure reliability. Candidates... 
    Flexible hours

    Akamai Technologies

    Cambridge, MA
    3 days ago
  • $75.2k - $95.3k

    WEX, Inc. is seeking a highly motivated entry-level Site Reliability Engineer (SRE) in Boston, MA. This role is integral to our SRE transformation, contributing to the performance and reliability of our systems. As an SRE, you'll monitor system reliability, manage incidents... 

    WEX, Inc.

    Boston, MA
    3 days ago
  • $127k - $249k

    The Team Platform Engineering is the department within SRE that is responsible for a range...  ...our developers to build and ship products to delight our customers. We manage the...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    1 day ago
  • $128k - $160k

     ...driver of progress, come build the future together. The Crown Is Yours As a Senior Site Reliability Engineer, you'll build and scale the critical infrastructure behind every product. In this role, you'll take on complex challenges across global data centers, multiple... 
    Full time
    Immediate start

    DraftKings Inc.

    Boston, MA
    2 days ago
  • $146.4k - $263.6k

     ...large datasets to analyze and measure the performance and reliability of our platform. We are networking data scientists: we...  ...project‑specific work. Driving partnership with Engineering, Operations and Product teams to help guide adoption of new features and processes... 
    Work at office

    Akamai Technologies

    Cambridge, MA
    2 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and...  ...our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    5 days ago
  • $150k

     ...daily at Coalition. About The Role We are looking for a Staff Site Reliability Engineer to lead AI enablement across our engineering...  ...to ensure AI‑generated output is reliable, secure, and production‑worthy. This role owns that layer. This role blends building... 
    Fixed term contract
    Work experience placement
    Work at office
    Remote work
    Home office
    Flexible hours
    Shift work

    Coalition, Inc.

    Boston, MA
    5 days ago
  • $127k - $249k

     ...are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain...  ...Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure...  ...everything we do results in a stronger product and a better experience for all Atlas... 
    Local area
    Remote work
    Flexible hours

    MongoDB

    Boston, MA
    4 days ago
  • $140k - $210.9k

    The Federal Reserve Bank of Boston seeks a Senior Site Reliability Engineer to enhance their payments service, FedNow. This position involves significant responsibilities in operating the production environment and the design of automation and monitoring tools to ensure... 
    Full time

    Dormont Manufacturing Co

    Boston, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer (Production Engineer). Be the first to apply!