Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$223k - $310k

Omnissa, LLC in

Sr. Manager, SRE & Performance (Finance) We are Omnissa ! Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from-anywhere experiences. We integrate industry-leading solutions-including Unified Endpoint Management, Virtual Apps and Desktops, Digital Employee Experience, and Security & Compliance-into a seamless, autonomous workspace that ada p ts to how people work. Our platform boosts employee engagement while optimizing IT operations, security, and cost. Guided by our Core Values- Act in Alignment, Build Trust, Foster Inclusiveness, Drive Efficiency, and Maximize Customer Value - we're growing rapidly and committed to delivering meaningful impact. If you're passionate about shaping the future of work, we'd love to hear from you. At Omnissa , we are committed to maintaining a fair, consistent, and secure hiring process for all candidates. As part of this approach, we use standard interview and verification practices designed to ensure alignment and protect both candidates and the organization. These practices are applied thoughtfully and with respect to candidate privacy. What is the opportunity? We are looking for a Senior Manager of Site Reliability Engineering (SRE) & Performance Engineering to lead teams responsible for the reliability, scalability, performance, and operational excellence of our distributed SaaS platform. This is a hands‑on technical leadership role for someone who can operate at the intersection of software engineering, distributed systems, cloud infrastructure, SRE, and performance engineering. You will lead engineers while partnering closely with development, architecture, security, SaaS operations, and product teams to ensure our services are designed and operated for scale, resilience, efficiency, and predictable performance. You will help establish engineering standards around SLOs, observability, performance testing, capacity planning, incident management, production readiness, and continuous reliability improvement. Here’s a breakdown: Lead and grow SRE and Performance Engineering teams, providing technical direction, coaching, career development, and establishing a strong culture of ownership and engineering excellence. Define and drive the organization's SRE and performance engineering strategy, including reliability goals, performance objectives, scalability standards, and operational readiness requirements. Partner with engineering teams to ensure systems are designed for high availability, scalability, fault tolerance, performance, and operability from the beginning rather than addressing these concerns after deployment. Establish and drive adoption of SLIs, SLOs, error budgets, service health indicators, and production readiness criteria for critical services. Lead performance engineering initiatives across distributed systems, including workload modeling, benchmarking, profiling, scalability testing, capacity planning, and bottleneck analysis across CPU, memory, I/O, storage, database, and network layers. Drive observability strategy across metrics, logs, traces, dashboards, and alerting, while improving signal quality and reducing noisy or non-actionable alerts. Provide technical leadership during complex production incidents, helping teams diagnose distributed system failures, latency spikes, resource contention, database issues, network problems, and cascading failures. Drive effective incident management, RCA/postmortem practices, corrective actions, and systemic reliability improvements, ensuring lessons from incidents translate into engineering changes. Partner with architects and senior engineers on system design and architecture reviews, challenging designs around scalability, resilience, failure modes, performance, data architecture, and operational complexity. Establish scalable approaches for performance regression detection and reliability validation within CI/CD pipelines, enabling issues to be identified earlier in the development lifecycle. Drive capacity management and forecasting, using production telemetry, workload characteristics, and performance models to anticipate infrastructure requirements and scalability limits. Improve platform efficiency through performance optimization, infrastructure right-sizing, and cost-aware engineering, balancing reliability, performance, and cloud cost. Champion automation and engineering-driven operations, reducing manual operational work and toil through software, tooling, Infrastructure as Code, and automated remediation. Establish measurable SRE and performance KPIs, such as availability/SLO attainment, MTTR, change failure rate, alert quality, performance regression rates, capacity headroom, operational toil, and recurring incident reduction. Collaborate across Product, Engineering, Security, Cloud/SaaS Operations, and Architecture organizations to deliver secure, resilient, scalable, and production-ready services. What will you bring to the company? Leadership & Engineering Management 12+ years of software engineering experience, with significant experience building and operating large-scale backend or distributed systems. 5+ years of engineering leadership/management experience, preferably leading SRE, Performance Engineering, Platform Engineering, Infrastructure, or backend engineering teams. Proven ability to build, mentor, and develop high-performing engineering teams, including senior and staff-level engineers. Strong ability to balance people leadership, technical strategy, operational priorities, and business objectives. Experience influencing engineering practices across teams without relying solely on organizational authority. Demonstrated ability to work effectively with senior engineers, architects, engineering managers, product leaders, security teams, and executive stakeholders. Technical Depth Strong understanding of distributed systems and microservices architectures, including scalability, availability, consistency, fault tolerance, failure modes, and architectural trade-offs. Strong software engineering background, preferably with experience in C#/.NET, Go or Java, and the ability to participate meaningfully in architecture, design, and code-level technical discussions. Deep understanding of Linux systems, including processes, memory, CPU scheduling, networking, file systems, containers, and system-level performance diagnostics. Strong experience with Docker and container orchestration using HashiStack technologies such as Nomad, Consul, and Vault. Experience operating high-scale data platforms using technologies such as Kafka, PostgreSQL, OpenSearch, Redis/Valkey, or equivalent technologies. Strong cloud experience, preferably with AWS, including services such as EC2, EKS, MSK, Aurora/RDS, OpenSearch, networking, storage, and cloud observability services. Strong understanding of CI/CD, Infrastructure as Code, deployment automation, release strategies, and modern DevOps practices. SRE & Performance Engineering Deep understanding of SRE principles, including SLIs/SLOs, error budgets, availability engineering, incident management, production readiness, operational toil, and reliability automation. Proven experience designing and implementing observability strategies using metrics, logging, tracing, dashboards, and actionable alerting. Strong understanding of performance engineering methodologies, including workload modeling, benchmarking, profiling, stress/load testing, scalability analysis, and performance regression detection. Experience diagnosing performance issues across applications, databases, operating systems, containers, infrastructure, and network layers. Experience with capacity planning and forecasting, including translating workload growth into infrastructure and service capacity requirements. Strong understanding of resilience engineering, including graceful degradation, retries, timeouts, circuit breakers, backpressure, rate limiting, disaster recovery, and failure testing. Demonstrated ability to turn production incidents and performance findings into systemic engineering improvements rather than tactical fixes. Location: Mountain View, CA Location Type: hybrid Education: Bachelor's Degree preferred, or equivalent combination of education and relevant professional experience. Compensation The typical base salary for this role is between USD $223,000 - $310,000 per year and it may be eligible for participation in a corporate bonus program. Actual compensation offer may vary from posted hiring range based upon geographic location, work experience, education, skill level, or other relevant factors. employee ownership health insurance 401k with matching contributions disability insurance paid-time off growth opportunities and more Omnissa is an Equal Employment Opportunity company and Prohibits Discrimination and Harassment of Any Kind: Omnissa is committed to the principle of equal employment opportunity and to providing a work environment free of discrimination and harassment. All employment decisions at Omnissa are based on business needs, job requirements and individual qualifications, without regard to race, color, religion, ancestry, ethnicity, national, social or ethnic origin, sex (including pregnancy), age, physical, mental or sensory disability, HIV status, sexual orientation, gender identity and/or expression, marital, civil union or domestic partnership status, past, present, or prospective service in the uniformed services, family medical history or genetic information, family or parental status, veteran status, or any other status protected by applicable laws or regulations in the locations where we operate. Omnissa will not tolerate discrimination or harassment based on any of these characteristics. Omnissa welcomes applicants of all ages. Omnissa will provide reasonable accommodations to applicants and employees who have protected disabilities consistent with applicable federal, state and local law. #J-18808-Ljbffr

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Mountain View, CA vacancy
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Suggested
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    1 day ago
  • $165k - $190k

     ...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-...  ...security platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering... 
    Suggested
    Work from home

    Obsidian Security

    Palo Alto, CA
    2 days ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are... 
    Suggested
    Remote work

    Abbott

    Sunnyvale, CA
    2 days ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is...  ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability... 
    Suggested
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    3 hours ago
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Full time

    Fiserv

    Sunnyvale, CA
    3 hours ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    1 day ago
  •  ...Google is seeking a Software Engineering Manager II in Site Reliability Engineering, based in Sunnyvale, California. This onsite role leads a team to ensure reliability and performance of critical systems, partnering with product and engineering teams to deliver scalable... 

    Epic Games

    Sunnyvale, CA
    5 days ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud... 
    Full time

    Saransh

    Sunnyvale, CA
    5 days ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data... 

    Tik Tok

    Mountain View, CA
    5 days ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $276.1k - $311.4k

     ...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    3 days ago
  • $145k - $165k

     ...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  •  ...Omnissa is seeking a Senior Manager of Site Reliability Engineering (SRE) & Performance to lead the reliability, scalability, and efficiency of our distributed SaaS platform. This hands-on technical leadership role operates at the intersection of software engineering,... 

    Omnissa, LLC in

    Mountain View, CA
    10 hours ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 

    Bolt Graphics, Inc.

    Sunnyvale, CA
    5 days ago
  •  ..., Elise AI, IBM and Accern. Position Summary We are hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep technical execution. You will be responsible for defining reliability... 
    Shift work

    Wand AI

    Palo Alto, CA
    5 days ago
  • $150.4k - $277.6k

     ...Technical Operations & Site Reliability Engineer, Customer SystemsAt Apple, Customer Experience is at the forefront of everything we do. The Customer Systems Operations team is looking for a highly skilled and motivated TechOps Engineer (Technical Operations & Site Reliability... 
    Work experience placement
    Relocation

    Apple

    Sunnyvale, CA
    1 day ago
  • $222k - $300.5k

     ...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps...  ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms... 

    JP Morgan Chase

    Palo Alto, CA
    3 hours ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 

    Google

    Mountain View, CA
    4 days ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    13 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $150k - $200k

     ...our CEO's funding announcement: The Reliability team owns the availability, performance,...  ...enforcing reliability standards across engineering Designing incident response processes...  ...ownership of production systems. As a Site Reliability Engineer on the Reliability... 
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    GrabJobs

    Sunnyvale, CA
    1 day ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...Science or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Mountain View, CA
    4 days ago
  • $262k - $364k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...degree in Computer Science or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Sunnyvale, CA
    2 days ago
  • $255.7k - $300k

    Lead a team of engineers to maintain service uptime while managing global on-call rotations...  ...improve operational practices to drive reliability, maintainability, and stakeholder alignment...  ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation... 
    Full time
    Work at office

    Google

    Sunnyvale, CA
    1 day ago
  • $172k - $300k

     ...Job Description GM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software. As one of our founding SREs... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $298k - $368k

     ...the Pipeline SRE Lead, you will play a key role in driving the reliability of our most critical release pipelines. You’ll lead the...  ...investigations and resolutions, as well as proactively partnering with engineering to evolve our software system to enhance their robustness and... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    4 days ago
  • $169k - $338k

     ...Regular/PermanentCompany: WalmartBusiness Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering organization, you will lead the technical development of next-generation agentic AI systems and... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    2 days ago
  •  ...function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!