Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Apex Systems Inc

Site Reliability Engineer


Location: Irvine, California (Working Onsite)

Role Overview

As a Site Reliability Engineer, you will join a team responsible for the reliability, scalability, and operational excellence of a product platform. This role involves working with engineering and product teams to design, build, and operate resilient platform services. A key focus will be building and owning a new custom observability platform to provide a clear, end-to-end view of critical pipelines. This position requires some coding and a willingness to learn.

Responsibilities:

Observability Platform Engineering

o Build and own a new observability platform that gives both engineers and business stakeholders a single, end-to-end view of the health of the company's critical multi-system pipelines.

o Bring together real-time telemetry for engineers and a curated, longer-term view for the business, so the organization can see not just why a request was slow but whether a pipeline as a whole is healthy and where it is failing.

o Map how work flows across systems so that failures can be traced to their source, downstream impact is understood, and every issue has a clear, accountable owner.

o Deliver the dashboards and views that let both technical and non-technical users understand pipeline health and drill into the underlying detail when needed.


Cloud & Infrastructure Monitoring

o Build capability for monitoring the health, performance, and availability of the company's multi-cloud infrastructure (AWS, GCP, and Azure) across enterprise-scale healthcare operations.

o Track the state of Kubernetes clusters and cloud-native services, surfacing scaling, resiliency, and reliability issues before they cause outages.

o Enhance observability for traffic and service-to-service communication across multi-cluster environments (for example, through the Istio service mesh) to catch latency and routing problems early.

o Provide visibility into environment consistency, disaster-recovery readiness, and rollout health so that availability and business-continuity goals are met for mission-critical applications.


Data Pipeline & Streaming Monitoring

o Develop monitoring capabilities for real-time stream processing (for example, Apache Flink) for throughput, latency, backpressure, and job failures.

o Track the health of event-streaming infrastructure such as Apache Kafka (Amazon MSK) and Kafka Connect, including consumer lag and data-ingestion issues.

o Enhance observability for change-data-capture flows (for example, Debezium and MongoDB) and schema management for breakages and compatibility problems.

o Assist with data orchestration team in enhancing observability for data-pipeline orchestration (for example, Dagster) so that failed or delayed jobs are detected and attributed quickly.


AI & ML Monitoring

o Track the performance, accuracy, and reliability of AI applications in production, including generative-AI and LLM/RAG workloads.

o Provide visibility into the AI-serving stack, such as model endpoints and vector databases, for latency, scaling, and indexing health.

o Develop AI guardrails and safety controls so that issues are detected and addressed quickly.


Security, Compliance & Observability

o Own the core monitoring and observability tooling (Prometheus, Grafana, OpenTelemetry) that provides distributed tracing and system-performance visibility.

o Maintain security controls such as automated credential rotation (Conjur Cloud) and default-deny network policies.

o Manage authentication and access for database clusters, including SASL/SSL (SCRAM-SHA-512) and RBAC permissions.

o Help keep infrastructure compliant with HIPAA, SOC 2, and ISO 27001 through automated policy enforcement and encrypted communications.

o Build "golden signal" dashboards (latency, traffic, errors, saturation) that populate automatically for every new microservice.

o Reduce operational toil through automation and self-healing remediation, and help track and optimize platform costs across all layers.


DevOps & Automation

o Maintain CI/CD pipelines (BitBucket Pipelines, ArgoCD) and GitOps-based automated testing and deployment.

o Support database deployment pipelines (Liquibase).

o Automate infrastructure and application configuration (Ansible).

o Enable automated health checks and failover for zero-downtime updates.

o Partner with and mentor engineering teams on reliability, observability, and DevOps best practices.


Qualifications:

Required:

• BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & education & experience in lieu of degree, 6+ years of relevant expertise is required.

• Strong, hands-on Kubernetes experience - deploying, scaling, and operating containerized workloads across multiple clusters and environments, including Helm-based packaging.

• Solid DevOps and GitOps practice, including CI/CD pipelines, infrastructure-as-code, and automated, GitOps-driven deployment (for example, with ArgoCD).

• Strong AWS skills, including object storage (S3), managed identity (IAM), secrets management, managed AI services, and managed Kubernetes (EKS).

• Working proficiency in React and TypeScript, sufficient to build and modify the platform's internal dashboard and graph-visualization UI.

• Strong Python engineering with a data-engineering focus, including building ETL services (incremental extraction, schema handling) and writing columnar data to object storage.


Preferred:

• BS degree in Computer Science or related field plus 4 years of relevant technology experience or equivalent combination of education & experience in lieu of degree, 6+ years of relevant expertise is required.

• Experience with graph databases, particularly Neo4j and Cypher, is a strong plus. The platform's curated business layer is built on Neo4j, so this experience is highly valued.

• Experience with Kafka real-time event streaming components and workloads.


Knowledge/Skills/Abilities:

• Ability to multi-task effectively without compromising the quality of the work.

• Excellent interpersonal, oral and written communication skills.

• Detail oriented, organized, process focused problem solver, proactive, ambitious, customer service focused.

• Ability to draw conclusions and make independent decisions with limited information.

• Ability to respond to common inquiries from customers, staff, regulatory agencies, vendors and other members of the business community.

• Self-motivated, reliable individual capable of working independently as well as part of the team.

• Motivated to drive improvement in a challenging environment.

• Strong background in data structures, algorithms and debugging.

• Demonstrated technical leadership, and successful participation in projects involving multiple engineers.

• Ability to learn quickly, understand complex systems and to work closely with others across multiple teams.

• Ability to handle uncertainty, time pressure and large technical challenges.

• Ability to deliver high-quality work on time.

• Strong attention to details, highly organized, computer literate


Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing® in Talent Satisfaction in the United States and Great Place to Work® in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click here for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Everforth Apex Benefits Overview: Everforth Apex offers a range of supplemental benefits, including medical, dental, vision, life, disability, and other insurance plans that offer an optional layer of financial protection. We offer an ESPP (employee stock purchase program) and a 401K program which allows you to contribute typically within 30 days of starting, with a company match after 12 months of tenure. Everforth Apex also offers a HSA (Health Savings Account on the HDHP plan), a SupportLinc Employee Assistance Program (EAP) with up to 8 free counseling sessions, a corporate discount savings program and other discounts. In terms of professional development, Everforth Apex hosts an on-demand training program, provides access to certification prep and a library of technical and leadership courses/books/seminars once you have 6+ months of tenure, and certification discounts and other perks to associations that include CompTIA and IIBA. Everforth Apex has a dedicated customer service team for our Consultants that can address questions around benefits and other resources, as well as a certified Career Coach. You can access a full list of our benefits, programs, support teams and resources within our 'Welcome Packet' as well, which an Everforth Apex team member can provide.

Everforth Apex Systems is an equal opportunity employer. We do not discriminate or allow discrimination on the basis of race, color, religion, creed, sex (including pregnancy, childbirth, breastfeeding, or related medical conditions), age, sexual orientation, gender identity, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, disability, status as a crime victim, protected veteran status, political affiliation, union membership, or any other characteristic protected by law. Everforth Apex will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable law.

If you require an accommodation under the Americans with Disabilities Act to participate in an interview with a virtual recruiter or to use our website for a search or application, please contact our Benefits Department at [email protected] or View phone number on click.appcast.io. Please note that this contact information is strictly to be used for medical ADA accommodations and that no other inquiries will be answered.

UnitedHealthcare creates and publishes the Transparency in Coverage Machine-Readable Files on behalf of Everforth Apex Systems.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Irvine, CA vacancy
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Suggested
    Full time
    Work at office

    Restaurant 365

    Irvine, CA
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Healthcare Payments team, you will solve complex and broad... 
    Suggested

    JP Morgan Chase

    Irvine, CA
    1 day ago
  • $166k - $250k

     ...to serve a wide variety of defense, IC and commercial customers in US and international markets.ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $166k - $220k

     ...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding...  ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $100k - $140k

     ...and software services. Our team of passionate engineers are constantly innovating, engineering...  ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced Site Reliability Engineer to join our team and play... 
    Local area
    Worldwide
    Weekend work

    TP-Link North America, Inc.

    Irvine, CA
    1 day ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 

    Chase

    Irvine, CA
    5 days ago
  • $166k - $220k

     ...Senior Site Reliability Engineer Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    anduril

    Costa Mesa, CA
    2 days ago
  •  ...Senior Site Reliability Engineer Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative... 

    anduril

    Irvine, CA
    2 days ago
  • $253k - $336k

     ...not years.ABOUT THE TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril's corporate systems.It drives...  ...its mission critical products.ABOUT THE JOB:The Director of Site Reliability Engineering owns the reliability system for the software that... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $191k - $253k

     ...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate...  ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril...  ...:10+ years of experience in site reliability engineering, production engineering... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $102.8k - $190.2k

    Team Name:Battle.net & Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436Job Description:This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering... 
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package

    Blizzard Entertainment

    Irvine, CA
    4 days ago
  • $145.7k - $218.5k

     ...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to...  ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated... 
    Work experience placement
    Shift work

    Sony Interactive Entertainment America

    Aliso Viejo, CA
    4 days ago
  • Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly...  ...for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability,... 
    Full time
    Temporary work

    Alcon Pharma

    Lake Forest, CA
    3 days ago
  • $146k - $194k

     ...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $158.6k - $237.6k

     ...End strategy and PD execution for a high-quality SS design delivery to SOC customers.What You Can ExpectThe IP Release Principal Engineer is a senior technical individual contributor within Marvell's IP Subsystem (IPSS) Center of Excellence (COE). This role is the quality... 
    Permanent employment
    Internship
    Work from home

    Marvell

    Irvine, CA
    3 days ago
  • $166k - $220k

     ...ABOUT THE JOB:We are seeking a highly skilled and driven Software Engineer to join a fast-paced team dedicated to the continued...  ...-quality code that enhances the performance, scalability, and reliability of the platformEngage in code reviews, architectural discussions... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $166k - $220k

     ...priorities, we want you to join Anduril’s Maritime Division and help us build the future of defense capability.About the JobSoftware Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation, data... 
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $37.26 - $68.93 per hour

    Team Name:DiabloJob Title:Software Engineer, Engine SystemsRequisition ID:R027570Job Description...  ...work week, with a mix of remote and on-site days. While hybrid is the standard...  ..., fixing bugs, and strengthening system reliability. Advance the game engine through performance... 
    Hourly pay
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package
    Flexible hours

    Blizzard Entertainment

    Irvine, CA
    12 hours ago
  • $166k - $220k

     ..., AI, computer vision, sensor fusion, and networking technology to the military in months, not years.WHY WE'RE HERE: The systems engineering team is looking for experienced systems engineers to support a Group 5 aircraft program at Anduril. The ideal candidate will have... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $166k - $220k

     ...in this contested warfighting domain.ABOUT THE JOBAs a Software Engineer, Systems Test for our Space team, you will be instrumental in...  ...threads. We work with mission partners and operators to deploy reliable and robust capabilities on operationally relevant fielding timelines... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $191k - $253k

     ...WHY WE ARE HERE We are seeking a software engineer for mission systems development of the Fury. Fury is a high-performance, multi-mission group 5 autonomous air vehicle (AAV) enabling trusted and collaborative autonomy for the high-end fight. Engineers in this role... 
    Full time
    Work experience placement
    Internship
    Relocation package

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $150k - $180k

    WHAT YOU’LL DOThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess a deep understanding of systems engineering and automation, including configuration management... 
    Work experience placement
    Local area

    Viant Technology

    Irvine, CA
    12 hours ago
  • $191k - $253k

     ...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $128k - $171k

     ...team builds the systems that keep our internal customers (the engineers, operators, and back-office teams across Anduril) productive....  ...processes end-to-end: turning manual, ticket-driven work into reliable software, applying AI where it makes sense, and shipping self-... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $102k - $173.4k

     ...Team: Advancing solution enablement, business practice maturity, and partner readiness. (IBM Software)As a Sr Technical Solutions Engineer, you need to be a technical leader and resource for SD&S and Ingram Micro. Maintain relevant applicable industry certifications and... 
    Full time
    Temporary work
    Work at office
    Remote work
    Worldwide
    Shift work

    Ingram Micro

    Irvine, CA
    2 days ago
  • $146k - $194k

     ...Reliability Engineer Costa Mesa, California, United States; Quincy, Massachusetts, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise... 
    Full time
    Work experience placement

    anduril

    Costa Mesa, CA
    2 days ago
  • $84.5k

     ...Facebook, Instagram ( , X ( and YouTube. ( Job Description The Reliability Engineer is responsible for driving continuous equipment and Utilities...  ...by managing the reliability improvement processes for the site, with the aim of reducing risk levels to the business. The... 
    Temporary work
    Work experience placement
    Local area

    AbbVie

    Irvine, CA
    2 days ago
  •  ...these assemblies and products throughout the qualification process ~ Advising and conferring with engineers in design review meetings to provide reliability findings and recommendations ~ Develop FMEA and FRACAS on MTS products. ~ Supports the... 

    CivicMinds, Inc

    Irvine, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!