Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$140k - $165k

instrumental-inc-

Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture digital exhaust and engineering context from assembly lines - images, test logs, BOM data, performance, repair cycles and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, Cisco, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production.

The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface - we believe we must have both the best technology and the best access to that technology to win.

As a Site Reliability Engineer, you'll operate, improve, and scale our AWS-based SaaS platform. You'll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational excellence. You'll participate in a bi-weekly on-call rotation, but the goal isn't simply to keep systems running - it's to continuously engineer away the operational complexity that comes with scaling our platform and customer base.

Requirements:
  • 3-4 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments.
  • Strong hands-on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3.
  • Experience managing infrastructure using Terraform or other Infrastructure as Code technologies.
  • Experience designing and supporting CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or similar platforms.
  • Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM.
  • Experience with Docker and Kubernetes.
  • Scripting experience with Python and/or Bash.
  • Experience supporting production environments through an on-call rotation, including incident response and root cause analysis.
  • Proven ability to take ownership of production issues and drive them through investigation, remediation, and long-term resolution.
Who You Are:
  • Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and continuously look for ways to make them more reliable, scalable, observable, and supportable.
  • Automation, automation, automation: If something is repetitive, manual, or error-prone, your first instinct is to automate it and make it disappear.
  • An engineer at heart: You don't want to repeatedly fight the same fires. You look for the underlying cause and build durable engineering solutions that reduce operational toil and technical debt.
  • Strong systems thinker: You understand how infrastructure, applications, networks, deployments, monitoring, and people interact—and can troubleshoot complex production issues across those boundaries.
  • Collaborative and reliable: You partner closely with software engineers to make services production-ready, improve operational workflows, and build reliability into systems before they become problems.
  • Comfortable with growth and ambiguity: You're comfortable making good decisions without perfect information and adapting as the platform, customer base, and company scale quickly.
Nice to Have:
  • Experience working in a high-growth B2B SaaS environment.
  • Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
  • Experience building internal tooling and automation to eliminate operational toil.
  • Experience supporting multi-region AWS environments.
  • AWS cost optimization or FinOps experience.
  • Network, application security, and compliance experience.
  • Experience introducing AI tools or processes into engineering and operational workflows.

This position requires access to items and data that are developed under U.S. government contracts and subject to dissemination controls that limit access to U.S. citizens only.

We're a growing team that works collaboratively, is supportive of each other, and is highly energized by the opportunity for a large impact. We actively work to promote an inclusive environment, valuing passion and the ability to learn.

The following is a representative annual base salary range for this position within the Bay Area: $140,000-$165,000. Job level and salary opportunities are evaluated through our interview process - we review the experience, knowledge, skills, and abilities of each applicant.

Instrumental is proud to offer a highly-rated variety of benefits, including health, vision, dental, commuter plans, and parental leave.

At Instrumental, protecting company and customer information is a shared responsibility. Employees are expected to comply with company engineering, security, access control, and privacy policies, and promptly report suspected security incidents or policy violations.

#J-18808-Ljbffr
Vacancy posted 21 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Palo Alto, CA vacancy
  •  ...Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes clusters (including GPU-backed clusters... 
    Suggested

    Amiri Recruiting

    Mountain View, CA
    2 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    2 days ago
  •  ...technologies. Our mission is to double America’s compute capacity without building new data centers. We are seeking a skilled Site Reliability Engineer to join our growing team. The ideal candidate will help ensure the reliability, scalability, and performance of our hybrid... 
    Suggested
    Work at office
    Weekend work

    FLUIX

    Palo Alto, CA
    2 days ago
  •  ...AI, IBM and Accern. Position Summary We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices... 
    Suggested
    Shift work

    Wand AI

    Palo Alto, CA
    2 days ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently... 
    Suggested
    Shift work

    AI Chopping Block

    Menlo Park, CA
    4 days ago
  • $150k - $180k

     ..., environmental, and innovation outcomes. Role Verrus is looking for candidates to serve as software-focused Senior Site Reliability Engineer at Verrus. This is a full‑time position based out of the Mountain View, CA office. Verrus takes a very technology‑forward... 
    Full time
    Work at office
    Local area
    Flexible hours

    Verrus, LLC

    Mountain View, CA
    2 days ago
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $170k - $200k

     ...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance... 
    Full time

    Zoomcar

    Sunnyvale, CA
    4 days ago
  • $132.6k - $214.5k

     ...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and...  ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    21 hours ago
  • $230k - $250k

     ...minds are shaping the future of network reliability, security, and AI‑ready operations. About...  ...you will be building the reliability engineering function at Forward — defining how we...  ...Looking For ~6+ years of experience in site reliability engineering, DevOps, or... 
    Night shift

    Forward

    Santa Clara, CA
    4 days ago
  •  ...Investigate and resolve performance and reliability issues across application, infrastructure, database, Kubernetes, and Linux layers...  ..., plan capacity, improve observability, and collaborate with engineering teams and business stakeholders. Requirements: Requires hands... 

    engineeringjobs.net, Inc.

    Sunnyvale, CA
    3 hours ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    21 hours ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  • $104.4k - $171k

     ...The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in... 

    Alibaba Cloud

    Sunnyvale, CA
    21 hours ago
  • $65 - $85 per hour

     ...Talent is partnering with Nvidia, a global leader in computer graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA. This is a full‑time (W‑2) contract role. Pay ranges from $65/hr to $85... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    21 hours ago
  • $150.4k - $277.6k

     ...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long...  ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced... 
    Relocation
    Day shift

    Apple

    Cupertino, CA
    1 day ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in... 
    Work experience placement

    Illumio

    Sunnyvale, CA
    2 days ago
  •  ...Site Reliability Engineer - 100% Remote Site Reliability Engineers (SREs) are responsible for working with different developer teams to keep our systems running smoothly. They are a blend of pragmatic operators and software craftspeople that apply excellent problem-... 
    Remote work
    Shift work

    Talentify.io

    Redwood City, CA
    1 day ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    4 days ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    a month ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    a month ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    21 hours ago
  • $120k - $150k

     ...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global...  ...every day? Join us at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    1 day ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, CA
    2 days ago
  • $262k - $364k

     ...and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely...  ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Sunnyvale, CA
    3 days ago
  • $175k - $265k

     ...Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-...  .... This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab,... 

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $195k - $285k

     ...purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the... 
    Remote work

    d-Matrix

    Santa Clara, CA
    21 hours ago
  •  ...Up to 25% Job ID: 1874 The Role The Platform Engineering team builds, secures, and operates scalable infrastructure...  ...products with on‑premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the platform stable and reliable... 
    Work at office
    Remote work

    Physics World

    Santa Clara, CA
    21 hours ago
  • $260k - $275k

     ...Saviynt Work on a mission-critical SaaS platform used by global enterprises Solve complex reliability challenges at scale Influence architecture and engineering culture at a company level Competitive compensation, benefits, and growth opportunities... 

    Saviynt

    Milpitas, CA
    21 hours ago
  • $152k - $287.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...such as Python, Go, Perl, or Ruby. ~ Mentored other engineers and influenced technical direction through design reviews, architecture... 
    Full time

    NVIDIA

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!