Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

AceStack LLC

Title: Site Reliability Engineer (SRE)


Location: Location: Sunnyvale, CA (3x/ week onsite)


Contract

Responsibilities:

  • Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions.
  • Operate, monitor, and triage all aspects of our production and non-production environments.
  • Collaborate on code, infrastructure, design reviews, and process enhancements Evaluate and integrate new technologies to improve system reliability, security, and performance.
  • Develop and implement automation to provision, configure, deploy, and monitor services.
  • Participate in an oncall rotation providing hands-on technical expertise during service impacting events.
  • Contribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.
Min Qualification:
  • 5+ years in Infrastructure Ops, Site Reliability Engineering, or DevOps focused role.
  • BS degree in computer science or equivalent field with 5+ years of experience.
  • Knowledge of Linux operating system principles, networking fundamentals, and systems management.
  • Demonstrable fluency in at least one of the following languages: Java, Python, or Go.
  • Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment.
  • Familiarity with micro-services architecture and container orchestration with Kubernetes.
  • Awareness of key security principles including encryption, keys (types and exchange protocols).
  • Understanding of SRE principals including monitoring, alerting, error budgets, fault analysis, and automation.
  • Strong sense of ownership, with a desire to communicate and collaborate with other engineers and teams.
  • Ability to identify and communicate technical and architectural problems, while working with partners and their team to iteratively find solutions.
Experience implementing automation


Scripting experience in Python


Enjoy building partnerships


Apple Exp is preferred

Role Descriptions: We are seeking a DevOps Site Reliability Engineer (SRE) with strong experience in containerization orchestration and automation. The ideal candidate will have hands-on expertise in Kubernetes Docker and Python and will be responsible for building scalable infrastructure automating operations and ensuring high availability of production systems.

Key Responsibilities
  • Design deploy and maintain containerized applications using Docker and Kubernetes.
  • Build and maintain automated infrastructure and deployment pipelines.Develop automation scripts and tools using Python.Manage and optimize Kubernetes clusters in production environments.
  • Implement CICD pipelines to streamline build| test| and deployment processes.
  • Monitor system performance and reliability using observability and monitoring tools.
  • Troubleshoot production issues and participate in incident response and root cause analysis.
  • Work closely with development teams to improve system reliability and deployment efficiency.
  • Implement security and best practices for container and cloud infrastructure.
Essential Skills:
  • We are seeking a DevOps Site Reliability Engineer (SRE) with strong experience in containerization orchestration and automation.
  • The ideal candidate will have hands-on expertise in Kubernetes Docker and Python and will be responsible for building scalable infrastructure automating operations and ensuring high availability of production systems.
Key Responsibilities
  • Design| deploy| and maintain containerized applications using Docker and Kubernetes.
  • Build and maintain automated infrastructure and deployment pipelines.
  • Develop automation scripts and tools using Python.
  • Manage and optimize Kubernetes clusters in production environments.
  • Implement CICD pipelines to streamline build test and deployment processes.
  • Monitor system performance and reliability using observability and monitoring tools.
  • Troubleshoot production issues and participate in incident response and root cause analysis.
  • Work closely with development teams to improve system reliability and deployment efficiency.
  • Implement security and best practices for container and cloud infrastructure.



Desirable Skills:


Skills: Digital : Cloud DevOps~Digital : Python~Digital : DevOps Continuous Integration and Continuous Delivery (CI/CD)~Digital : Kubernetes~Digital : Site Reliability Engineering (SRE) Experience
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Sunnyvale, CA vacancy
  •  ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python... 
    Suggested

    Purple Drive

    Sunnyvale, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job Description Deploy, support and monitor new and existing services, platforms, and application stacks. Use scale testing to measure, tune... 
    Suggested

    Zortech Solutions

    Cupertino, CA
    2 days ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS)... 
    Suggested
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    1 day ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    3 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Suggested
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    3 days ago
  • $175k - $229k

     ...DevOps Engineer Instrumental builds the manufacturing acceleration platform behind the...  ...Requirements: ~5 or more years of DevOps or SRE experience deploying and operating...  ...KPIs to ensure ongoing performance, reliability and efficiency. ~ Network/application... 

    Instrumental Inc

    Palo Alto, CA
    2 days ago
  •  ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. • Experience with web applications and distributed systems... 

    Omega Solutions

    Santa Clara, CA
    1 day ago
  •  ...SRE Engineer Location: Sunnyvale CA Rate: DOE Duration: 12+ Months What You Will Do: Identify, develop and execute opportunities to raise the bar on engineering & operational excellence. Providing thought leadership and defining strategy on developer... 

    Redolent

    Sunnyvale, CA
    1 day ago
  •  ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote... 
    Full time
    Local area
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    1 day ago
  • $132.6k - $214.5k

     ..., you will collaborate closely with our engineering teams to develop innovative solutions that...  ...and health. As a Senior Staff SRE with the Cortex Observability team, you...  ...operability of the product and ensure the reliability and availability of our services. Qualifications... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $230k - $250k

     ...shaping the future of network reliability, security, and AI‑ready...  ...is not a "keep the lights on" SRE role. As our first or early SRE...  ...be building the reliability engineering function at Forward — defining...  ...~6+ years of experience in site reliability engineering, DevOps... 
    Night shift

    Forward

    Santa Clara, CA
    14 hours ago
  • $170k - $200k

     ...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining...  ...background in infrastructure automation, system reliability, and a SRE mindset of continuous improvement. Key Responsibilities... 
    Full time

    Zoomcar

    Sunnyvale, CA
    15 hours ago
  • $150.4k - $277.6k

     ...States Software and Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple...  ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role... 
    Relocation
    Day shift

    Apple

    Cupertino, CA
    2 days ago
  • $145k - $165k

     ...Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  •  ...Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our...  ...role — no customer interaction. Must-Have: ~4+ years in SRE, DevOps, or Infrastructure Engineering ~ Solid experience... 

    Amiri Recruiting

    Mountain View, CA
    3 days ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the impact...  ...have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    1 day ago
  •  ...world running. Location: 5 On-Site Days a Week in Sunnyvale, CA Headquarters Our Engineering team is driven by a culture...  ...history. Your Impact As an SRE Engineer II, you will be responsible...  ...will work on enhancing system reliability and scalability of Illumio SaaS... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  • $81.5k - $141.3k

     ...Site Reliability Engineer II Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life...  ...skilled and mission-driven Site Reliability Engineer (SRE) to join our DevOps team. In this critical role, you will be... 
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  • $110k - $130k

     ...with the World's leading AI-first Quality Engineering Company? Ready to advance your career,...  ...at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in...  ...rotation and support production Incidents. SRE Skillsets - Expectations from Pricing &... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    2 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Contract work

    VDart

    Santa Clara, CA
    21 hours ago
  •  ...the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...for an experienced Senior Site Reliability Engineer (SRE) with a strong background in AWS & Azure... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  • $152k - $287.5k

     ...infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked hands-...  ...through repair. ~ Experience managing production reliability through on-call duties, incident response, observability, and... 
    Permanent employment
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD Roles & Responsibilities • Extensive experience working with linux flavors like rhel/centos os, shells, filesystems and utilities • Knowledge of distributed computing and experience... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    1 day ago
  •  ...Role : SRE Engineer Location : San Jose, CA (ONSITE) FULL TIME ONLY Job Description Must Have Technical/Functional Skills: • Exp. in Apache SPARK development, Kubenetes, CI-CD Pipeline, Jenkins, Dockers, Kubernetes, PL SQL, Python... 
    Full time

    AceStack LLC

    San Jose, CA
    3 days ago
  • Coding experience in one or more of Python or Java. Experienced with automating infrastructure with scripting (Shell Script, Python) and tooling (Puppet, Terraform, Ansible, Chef, etc). Experienced with Splunk for investigating or monitoring problems on systems...
    Flexible hours

    VBeyond

    Sunnyvale, CA
    3 days ago
  • • Design, implement, and maintain complex data systems supporting millions of customers with Cloud Native principles and best practices to ensure highly available, secure, performant and scalable database systems • Build and maintain CI/CD pipelines in Jenkins • Build...

    United IT Solutions

    Mountain View, CA
    21 hours ago
  •  ...ability to multi-task in fast-paced environments. Global Collaboration: Comfortable working cross-functionally with product/engineering units across multiple time zones. Documentation: High care in creating detailed design specifications and presenting... 

    VBeyond

    Santa Clara, CA
    1 day ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...Responsibilities Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching, and career development...  .... Preferred Qualifications Experience managing SRE, DevOps, Production Engineering, or infrastructure teams... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    15 hours ago
  • $200k - $260k

     ...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical...  ...and eliminating work through automation. On the SRE team, you'll have the opportunity to manage the complex challenges... 
    Work at office
    Home office

    Glean - Mountain View, CA, US

    Mountain View, CA
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!