Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer (SRE)

$175k - $229k

Instrumental Inc

Senior Site Reliability Engineer (SRE)

Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, L3Harris, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production. The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface – we believe we must have both the best technology and the best access to that technology to win.

Requirements:
  • 5 or more years of DevOps or SRE experience deploying and operating commercial SaaS platforms on public cloud infrastructure, AWS preferred.
  • Expert knowledge with Linux, shell, containerization, Kubernetes, IaC (terraform preferred), monitoring, logging, and APM tools.
  • Proven ability to take initiative and drive impactful projects to completion efficiently and independently.
  • Comfort with ambiguity, pace, and frequent pivots inherent in a startup environment, with a track record of creating clarity for teams.
  • Experience introducing and integrating AI tools/processes into development and operation workflows.
  • Demonstrated skill in setting, iterating on, and measuring KPIs to ensure ongoing performance, reliability and efficiency.
  • Network/application security and compliance experience is a plus.
Who You Are:
  • Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and sweat the details around latency, uptime, and scale.
  • Systems engineering & infrastructure expertise: You've spent real time building and running distributed systems and know your way around cloud infrastructure, networks, and operating systems.
  • Automation, automation, automation: If something is repetitive or error-prone, your first instinct is to automate it and make it disappear.
  • Operating in ambiguity & high-growth environments: You're comfortable making good calls without perfect information and adapting as the system and company grow fast.
  • Dependable, trustworthy: People trust you to own problems, show up when things are broken, and follow through.

This position requires access to items and data that are developed under U.S. government contracts and subject to dissemination controls that limit access to U.S. citizens only. We're a growing team that works collaboratively, is supportive of each other, and is highly energized by the opportunity for a large impact. We actively work to promote an inclusive environment, valuing passion and the ability to learn. You're encouraged to apply even if your experience doesn't precisely match the job description! The following is a representative annual base salary range for this position within the Bay Area: $175-229k. We consider candidates at multiple levels for this role. Job level and salary opportunities are evaluated through our interview process – we review the experience, knowledge, skills, and abilities of each applicant. Instrumental is proud to offer a highly-rated variety of benefits, including health, vision, dental, commuter plans, and parental leave. At Instrumental, protecting company and customer information is a shared responsibility. Employees are expected to comply with company engineering, security, access control, and privacy policies, and promptly report suspected security incidents or policy violations.

Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer (SRE) in Palo Alto, CA vacancy
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...as-a-Service (CVaaS) global SRE team. SREs at Arista combine...  ...EngineeringExperience level: Mid-Senior LevelIndustry: Computer... 
    Senior

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $186.9k - $267.7k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San...  ...AI agents behave as intended, improving reliability and reducing risks. This unified approach...  ...and control.As a Staff Site Reliability Engineer (SRE), you will provide technical leadership... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Palo Alto, CA
    15 hours ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    19 hours ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Suggested
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    4 days ago
  • $210.6k - $305.1k

     ...soil. Lead, inspire, and develop a talented SRE team, fostering a culture of innovation,...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...open-source projects, particularly in SRE, observability, or AI/ML domains, and... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    19 hours ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems...  ...Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    4 days ago
  • $137.77k - $194.59k

     ...of roughly 80 scientists and engineers building and operating Rubin'...  ...role: \n You will own the reliability and robustness of Rubin...  ...Experience working in an SRE, DevOps, or data-intensive systems...  ...position, SLAC is open to on-site, hybrid, and remote work options... 
    Senior
    Remote work
    Flexible hours
    Night shift

    Stanford University

    Menlo Park, CA
    4 days ago
  •  ...this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by...  ...position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...-source projects, particularly in SRE, observability, or AI/ML domains, and... 
    Senior

    J.P. Morgan

    Palo Alto, CA
    2 days ago
  • $207k - $301k

     ...influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on...  ...of complex distributed systemsSite Reliability Engineering (SRE) combines software and systems engineering to build and run... 

    Google

    Sunnyvale, CA
    3 days ago
  •  ...Title: Site Reliability Engineer (SRE) Location: Location: Sunnyvale, CA (3x/ week onsite) Contract Responsibilities: Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure... 
    Contract work

    AceStack LLC

    Sunnyvale, CA
    2 days ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS)... 
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    4 days ago
  •  ...automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud...  ...Who You AreExperienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    19 hours ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...operating production distributed systems as SRE/DevOps/Platform Ops.Proven ownership of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $90k - $180k

     ...serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,...  ...and mission-driven Senior Site Reliability Engineer (SRE) to join our DevOps team. In this critical role, you will... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and help...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $232k - $263k

     ...growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian, you will define and...  ...strategic partner to DevOps and Platform Engineering leadership, shaping a unified...  ...roles3+ years operating at a senior or technical leadership level (Staff... 
    Senior
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    19 hours ago
  • $174k - $253k

     ...QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5...  ...degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations as if it’s a... 
    Senior

    Socket

    Sunnyvale, CA
    4 days ago
  •  ...Senior Sre For Gpu Infrastructure You'll own the GPU infrastructure Luma's research and...  ...you keep training and inference clusters reliable and fast, and you help redesign them for...  ...metal role for a first-principles Linux engineer. You'll be the final escalation for the... 
    Senior
    Work experience placement

    Luma AI

    Redwood City, CA
    3 days ago
  • $150k - $175k

     ...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed... 
    Senior
    Remote work

    ASAPP

    Mountain View, CA
    19 hours ago
  •  ...Senior Site Reliability Engineer Latitude AI is building the future of Ford's autonomy roadmap to make travel safer, less stressful, and more enjoyable for everyone. Bringing this vision to scale, our fully in-house developed hands-free ADAS platform will debut on... 
    Senior
    Work at office
    Immediate start

    Latitude AI

    Palo Alto, CA
    4 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer...  ...compliance, and uptime requirements. ~ Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident... 
    Senior

    Kody

    Palo Alto, CA
    17 days ago
  • $165k - $190k

     ...term growth and IPO readiness.About the DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    2 days ago
  • $120k - $180k

     ...The future of cybersecurity starts with you.Sr. SRE & DevOps EngineerAbout the Role:At CrowdStrike, our engineering organization depends on shared infrastructure platforms...  ...dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature... 
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    4 days ago
  • $144k - $230k

     ...join the team and see how you can make a lasting impact on the world.We are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. Applicants with SRE or equivalent experience are encouraged.What you will be doing:You will build... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. • Experience with web applications and distributed systems... 

    Omega Solutions

    Santa Clara, CA
    4 days ago
  •  ...Acryl Data seeks a Site Reliability Engineering (SRE) Tech Lead to enhance the reliability and scalability of its DataHub platform. The role involves leading infrastructure design, optimizing system performance, and driving continuous improvement across cloud deployments... 

    Acryl Data

    Palo Alto, CA
    19 hours ago
  •  ...Google, Amazon, Miro, Elise AI, IBM and Accern. Position Summary We are hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep technical execution. You will be... 
    Shift work

    Wand AI

    Palo Alto, CA
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer (SRE). Be the first to apply!