Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

James Consultancy Services , LLC

Job Description

Job Description

Role Description

We’re seeking an experienced, highly collaborative SRE to partner with product teams

and tackle our most critical infrastructure challenges. You’ll be hands‑on in designing,

building, and operating our cloud platform—and driving the reliability, performance, and

security that empower our engineering organization.

As a Site Reliability Engineer at DrumWave, you will:

  • Infrastructure as Code & CI/CD: Automate provisioning and deployments with Terraform and integrate best-practice pipelines (GitHub Actions, ArgoCD, etc.).
  • Reliability Engineering: Define SLIs/SLOs, manage error budgets, and build dashboards & alerts to proactively measure and improve system health.
  • Security & Compliance: Enforce least-privilege IAM policies, automate vulnerability scans, and maintain audit logging for compliance.
  • Monitoring & Observability: Instrument services with metrics, logs, and distributed tracing to enable rapid troubleshooting, aid teams in alerting, custom metrics, and dashboarding.
  • Incident Management: Own on-call rotations, lead real-time incident response, conduct post-mortems, and drive continuous improvements.
  • Cost Optimization: Implement tagging strategies, right-size resources, and leverage concrete data to decide on optimal methods to control cloud spend at scale.
  • Documentation & Mentorship: Author runbooks, standards, and best-practice guides—and coach dev teams on implementing modern DevOps, reliability, and security patterns.

Qualifications

  • Have 5+ years of experience running production critical systems
  • Deep proficiency with the AWS Cloud and Cloud-Native best practices
  • Experience with Kubernetes (EKS, GKE) and Container Orchestration at scale
  • Skilled in Terraform to declaratively provision and maintain infrastructure services
  • Working knowledge of managing and debugging databases like Redis and PostgresStrong familiarity with VPC, VPN, Load Balancing, and cloud networking components
  • Proficiency with Git workflows, branching strategies, and CI/CD system integrations
  • Solid understanding of web and network protocols and standards ( REST, TLS, DNS, etc...)

Nice to Have’s:

  • Bachelor's degree, or equivalent in Computer Science, Engineering, or a related field.
  • Experience with ArgoCD, Github Actions, Jenkins, or other CI/CD pipeline solutions
  • Working knowledge of Python, Golang, and Helm templating languages
  • Node.js experience a plus, including running scalable, resilient Node microservices
  • Grasp of foundational security best practices for cloud infrastructure
  • Awareness of Terragrunt, managing Terraform state, and optimal project structure
  • Seasoned in production readiness fundamentals amidst a fast moving team

Industry

  • Technology, Information and Internet
#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Mountain View, CA vacancy
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Suggested
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    2 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Suggested
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    3 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    2 days ago
  • $104.4k - $171k

     ...The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in... 
    Suggested

    Alibaba Cloud

    Sunnyvale, CA
    1 day ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in... 
    Work experience placement

    Illumio

    Sunnyvale, CA
    2 days ago
  •  ...Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes clusters (including GPU-backed clusters... 

    Amiri Recruiting

    Mountain View, CA
    2 days ago
  • $140k - $165k

     ...'s most complex electronics. We capture digital exhaust and engineering context from assembly lines - images, test logs, BOM data, performance...  ...and the best access to that technology to win. As a Site Reliability Engineer, you'll operate, improve, and scale our AWS-based... 

    instrumental-inc-

    Palo Alto, CA
    1 day ago
  •  ...Investigate and resolve performance and reliability issues across application, infrastructure, database, Kubernetes, and Linux layers...  ..., plan capacity, improve observability, and collaborate with engineering teams and business stakeholders. Requirements: Requires hands... 

    engineeringjobs.net, Inc.

    Sunnyvale, CA
    11 hours ago
  •  ...that keep the world running. Location: 5 On-Site Days a Week in Sunnyvale, CA Headquarters Our Engineering team is driven by a culture that thrives on visionary...  ...to-day basis, you will work on enhancing system reliability and scalability of Illumio SaaS products, and... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    1 day ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    1 day ago
  •  ...technologies. Our mission is to double America’s compute capacity without building new data centers. We are seeking a skilled Site Reliability Engineer to join our growing team. The ideal candidate will help ensure the reliability, scalability, and performance of our hybrid... 
    Work at office
    Weekend work

    FLUIX

    Palo Alto, CA
    2 days ago
  • $129k - $158k

     ...Site Reliability Engineer Join Fortinet, a cybersecurity pioneer with over two decades of excellence, as we continue to shape the future of cybersecurity and redefine the intersection of networking and security. At Fortinet, our mission is to safeguard people, devices... 
    Full time
    Worldwide

    Edelman

    Sunnyvale, CA
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Work at office

    Chase

    Palo Alto, CA
    11 hours ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    11 hours ago
  • $139.73k - $160.37k

     ...job description explicitly states otherwise, all roles are on-site five days per week at one of our offices in McLean, VA;...  ...which can be found here. Role Overview We are seeking a Site Reliability Engineer to join our Core Platform Engineering organization. The SRE team... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    ID.me

    Mountain View, CA
    3 days ago
  • $81.5k - $141.3k

     ...Site Reliability Engineer II Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices... 
    Remote work

    Abbott

    Sunnyvale, CA
    11 hours ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, CA
    2 days ago
  • $222k - $300.5k

     ...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps...  ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    2 days ago
  •  ...AI, IBM and Accern. Position Summary We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices... 
    Shift work

    Wand AI

    Palo Alto, CA
    2 days ago
  • $200k - $260k

     ...for enterprise trust, as we bring Work AI to every employee, in every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing,... 
    Work at office
    Home office
    Flexible hours

    Glean.info

    Mountain View, CA
    2 days ago
  • $150k - $180k

     ..., environmental, and innovation outcomes. Role Verrus is looking for candidates to serve as software-focused Senior Site Reliability Engineer at Verrus. This is a full‑time position based out of the Mountain View, CA office. Verrus takes a very technology‑forward... 
    Full time
    Work at office
    Local area
    Flexible hours

    Verrus, LLC

    Mountain View, CA
    2 days ago
  • $169k - $338k

     ...advanced agentic AI systems that can autonomously handle complex reliability engineering workflows, predictive failure analysis, and self-...  ...associates, or business operations across any Walmart system.Site Reliability Engineering Technical Excellence:Design, write and... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    3 days ago
  •  ...function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.As a... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $104.9k - $174.7k

     ...SRE role is responsible for improving the reliability, availability, performance, and...  ...actions through completion.Follow up with engineering, development, security, support, and business...  ...Qualifications5+ years of experience in Site Reliability Engineering, Systems Engineering... 
    Full time
    Local area

    LexisNexis Risk Solutions Group

    San Jose, CA
    4 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!