Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

AceStack LLC

Title: Site Reliability Engineer (SRE)


Location: Location: Sunnyvale, CA (3x/ week onsite)


Contract

Responsibilities:

  • Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions.
  • Operate, monitor, and triage all aspects of our production and non-production environments.
  • Collaborate on code, infrastructure, design reviews, and process enhancements Evaluate and integrate new technologies to improve system reliability, security, and performance.
  • Develop and implement automation to provision, configure, deploy, and monitor services.
  • Participate in an oncall rotation providing hands-on technical expertise during service impacting events.
  • Contribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.
Min Qualification:
  • 5+ years in Infrastructure Ops, Site Reliability Engineering, or DevOps focused role.
  • BS degree in computer science or equivalent field with 5+ years of experience.
  • Knowledge of Linux operating system principles, networking fundamentals, and systems management.
  • Demonstrable fluency in at least one of the following languages: Java, Python, or Go.
  • Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment.
  • Familiarity with micro-services architecture and container orchestration with Kubernetes.
  • Awareness of key security principles including encryption, keys (types and exchange protocols).
  • Understanding of SRE principals including monitoring, alerting, error budgets, fault analysis, and automation.
  • Strong sense of ownership, with a desire to communicate and collaborate with other engineers and teams.
  • Ability to identify and communicate technical and architectural problems, while working with partners and their team to iteratively find solutions.
Experience implementing automation


Scripting experience in Python


Enjoy building partnerships


Apple Exp is preferred

Role Descriptions: We are seeking a DevOps Site Reliability Engineer (SRE) with strong experience in containerization orchestration and automation. The ideal candidate will have hands-on expertise in Kubernetes Docker and Python and will be responsible for building scalable infrastructure automating operations and ensuring high availability of production systems.

Key Responsibilities
  • Design deploy and maintain containerized applications using Docker and Kubernetes.
  • Build and maintain automated infrastructure and deployment pipelines.Develop automation scripts and tools using Python.Manage and optimize Kubernetes clusters in production environments.
  • Implement CICD pipelines to streamline build| test| and deployment processes.
  • Monitor system performance and reliability using observability and monitoring tools.
  • Troubleshoot production issues and participate in incident response and root cause analysis.
  • Work closely with development teams to improve system reliability and deployment efficiency.
  • Implement security and best practices for container and cloud infrastructure.
Essential Skills:
  • We are seeking a DevOps Site Reliability Engineer (SRE) with strong experience in containerization orchestration and automation.
  • The ideal candidate will have hands-on expertise in Kubernetes Docker and Python and will be responsible for building scalable infrastructure automating operations and ensuring high availability of production systems.
Key Responsibilities
  • Design| deploy| and maintain containerized applications using Docker and Kubernetes.
  • Build and maintain automated infrastructure and deployment pipelines.
  • Develop automation scripts and tools using Python.
  • Manage and optimize Kubernetes clusters in production environments.
  • Implement CICD pipelines to streamline build test and deployment processes.
  • Monitor system performance and reliability using observability and monitoring tools.
  • Troubleshoot production issues and participate in incident response and root cause analysis.
  • Work closely with development teams to improve system reliability and deployment efficiency.
  • Implement security and best practices for container and cloud infrastructure.



Desirable Skills:


Skills: Digital : Cloud DevOps~Digital : Python~Digital : DevOps Continuous Integration and Continuous Delivery (CI/CD)~Digital : Kubernetes~Digital : Site Reliability Engineering (SRE) Experience
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Sunnyvale, CA vacancy
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software... 
    Suggested

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $207k - $301k

     ...influential relationships with multiple stakeholders across the Site Reliability Engineering and Developer organizations.Serve as an expert on...  ...of complex distributed systemsSite Reliability Engineering (SRE) combines software and systems engineering to build and run... 
    Suggested

    Google

    Sunnyvale, CA
    4 days ago
  • $186.9k - $267.7k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San...  ...AI agents behave as intended, improving reliability and reducing risks. This unified approach...  ...and control.As a Staff Site Reliability Engineer (SRE), you will provide technical leadership... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    Sunnyvale, CA
    22 hours ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS)... 
    Suggested
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    5 days ago
  • $120k - $180k

     ...The future of cybersecurity starts with you.Sr. SRE & DevOps EngineerAbout the Role:At CrowdStrike, our engineering organization depends on shared infrastructure platforms...  ...dedicated engineering ownership to operate reliably, scale safely, harden for security, and mature... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    5 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    2 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    1 day ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    5 days ago
  • $175k - $229k

     ...Senior Site Reliability Engineer (SRE) Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles... 

    Instrumental Inc

    Palo Alto, CA
    12 hours ago
  •  ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. • Experience with web applications and distributed systems... 

    Omega Solutions

    Santa Clara, CA
    5 days ago
  •  ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote... 
    Full time
    Local area
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    5 days ago
  • $60 - $64 per hour

     ...Request ID: 100379-1 Title: DevOps / SRE Engineer Location : Sunnyvale C Duration: 6+ Months Salary Range: $60- $64 an hour on W2 or C2C Job Description: Role Descriptions: • Support and administer large-scale Kubernetes platforms... 

    Artech

    Sunnyvale, CA
    3 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining,...  ...background in infrastructure automation, system reliability, and a SRE mindset of continuous improvement.Key Responsibilities:... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    3 days ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the impact...  ...have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will... 
    Night shift

    Forward Networks

    Santa Clara, CA
    1 day ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...operating production distributed systems as SRE/DevOps/Platform Ops.Proven ownership of... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure....  ...Who You AreExperienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute Farm team and help build...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-...  ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through... 
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $90k - $180k

     ...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA...  ...and mission-driven Senior Site Reliability Engineer (SRE) to join our DevOps team. In this critical role, you will be... 
    Remote work

    Abbott

    Sunnyvale, CA
    3 days ago
  • $145k - $165k

     ...Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    2 days ago
  • $64 - $68 per hour

     ...Akkodis is seeking a Site Reliability Engineer for a Contract with a client in Sunnyvale, CA/Austin, TX (Hybrid). The ideal candidate with...  ...~6-8 years of experience in Site Reliability Engineering (SRE), Cloud Operations, DevOps, or Infrastructure Engineering.... 
    Hourly pay
    Contract work
    Temporary work
    Local area

    Akkodis

    Sunnyvale, CA
    2 days ago
  • $150k - $195k

     ...team is growing, and we are looking for engineers with passion for automation. You will help...  ...teams to improve the scalability and reliability of internal processes. Participate in an...  ...Minimum Qualifications 3 years of Devops/SRE experience with production systems (depending... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud... 
    Full time

    Saransh

    Sunnyvale, CA
    1 day ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  •  ...world running. Location: 5 On-Site Days a Week in Sunnyvale, CA Headquarters Our Engineering team is driven by a culture...  ...history. Your Impact As an SRE Engineer II, you will be responsible...  ...will work on enhancing system reliability and scalability of Illumio SaaS... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    2 days ago
  • $210.6k - $305.1k

     ...soil. Lead, inspire, and develop a talented SRE team, fostering a culture of innovation,...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Los Altos, CA
    3 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency...  ...service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174k - $253k

     ...QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5...  ...degree in Computer Science or Engineering. ABOUT THE JOB: Site Reliability Engineering (SRE) is what you get when you treat operations as if it’s a... 

    Socket

    Sunnyvale, CA
    5 days ago
  •  ...ability to multi-task in fast-paced environments. Global Collaboration: Comfortable working cross-functionally with product/engineering units across multiple time zones. Documentation: High care in creating detailed design specifications and presenting... 

    VBeyond

    Santa Clara, CA
    5 days ago
  •  ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD Roles & Responsibilities • Extensive experience working with linux flavors like rhel/centos os, shells, filesystems and utilities • Knowledge of distributed computing and experience... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    5 days ago
  •  ...Role : SRE Engineer Location : San Jose, CA (ONSITE) FULL TIME ONLY Job Description Must Have Technical/Functional Skills: • Exp. in Apache SPARK development, Kubenetes, CI-CD Pipeline, Jenkins, Dockers, Kubernetes, PL SQL, Python... 
    Full time

    AceStack LLC

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!