Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Staff Site Reliability Engineer

$207.4k - $259.2k

Archer Aviation

Headquartered in Silicon Valley, California, Archer is a leader in the next-gen aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”), aviation-related physical artificial intelligence (“AI”) solutions, and other technologies to customers worldwide across the commercial aerospace and defense sectors.Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability, performance, and security of our core systems and services. You will leverage your extensive expertise in various technologies to design, implement, and maintain robust infrastructure and automation solutions.ResponsibilitiesImplement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS).Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems.Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms.Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management.Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations.Implement and enforce robust security practices across our infrastructure, applications, and data.Design and maintain Docker-based containerization solutions for our applications.Develop and maintain automation scripts and tools using Python, Bash, and PowerShell.Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception.Troubleshoot complex production issues across various layers of the stack, identifying root causes and implementing preventative measures.Participate in on-call rotations to support production systems.Qualifications12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence.Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting.Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar.Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark).Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD).Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles.Comprehensive understanding of security best practices for cloud environments, applications, and data.Proficiency in Docker for containerization and orchestration.Advanced scripting and programming skills in Python, Bash, and PowerShell.Solid understanding of networking concepts, distributed systems, and operating systems.Excellent problem-solving, analytical, and communication skills.Ability to work independently and as part of a highly collaborative team.Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.Preferred QualificationsExperience with other Kubernetes distributions or cloud providers.Familiarity with compliance frameworks (e.g., SOC 2, HIPAA, GDPR).Certifications in AWS, Kubernetes, or other relevant technologies.Successful candidates must be able to demonstrate U.S. citizenship, permanent residency, or status as a protected individual to satisfy ITAR, contractual, and/or regulatory requirements.Please note that this job description is intended to provide a general overview of the position and does not include an exhaustive list of responsibilities and qualificationsAt Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company’s business strategy. For this position we are targeting a base pay between $207,400 - $259,200. Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience.Archer is proud to be an Equal Opportunity employer committed to diversity and inclusivity in the workplace. All aspects of employment are decided on the basis of merit, qualifications, and business needs. We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state or local laws.Archer is committed to working with and providing reasonable accommodations to job applicants with physical or mental disabilities, and those with sincerely held religious beliefs. Applicants who may require reasonable accommodation for any part of the application or hiring process should provide their name and contact information to Archer’s People Team at View email address on us.fitly.work. Reasonable accommodations will be determined on a case-by-case basis.Information collected and processed as part of any job applications you choose to submit is subject to Archer's Candidate Privacy Policy.Certain positions may be eligible for visa sponsorship.Archer is proud to be an Equal Opportunity employer committed to diversity and inclusivity in the workplace. All aspects of employment are decided on the basis of merit, qualifications, and business needs. We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state or local laws.Archer Aviation does not engage with external recruiting agencies/individual recruiters with whom it does not have a prior written agreement. Archer reserves the right to make use of any unsolicited resumes that it receives and bears no responsibility for payment of any fees asserted from the use of unsolicited resumes. If you are a recruiting agency or individual recruiter wishing to do business with Archer, please reach out to View email address on us.fitly.work. All employment processes are managed by the Archer People Team.

Vacancy posted 9 days ago
Similar jobs that could be interesting for youBased on the Sr Staff Site Reliability Engineer in San Jose, CA vacancy
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,...  ...virtualization technologies, SR-IOV, and DPDKUnderstanding of... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  •  ...you want to shape the physical hardware that frontier models train on, this is the role.We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive... 
    Senior

    Amazon

    Cupertino, CA
    3 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $267k - $356k

     ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...in the industry, which means reliability and performance aren't just...  ...across new and existing sites using tools such as Ansible,...  ...CSI drivers.Experience with SR-IOV and virtualization (KVM/QEMU... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  • $160k - $190k

     ...an embedded semiconductor solution provider driven by its Purpose, To Make Our Lives Easier . With a global team of over 21,000 engineers and problem solvers in more than 30 countries, we offer the opportunity to work on world‑leading technology for Automotive, Industrial... 
    Senior
    Permanent employment
    Work experience placement
    Remote work

    Renesas Electronics

    San Jose, CA
    4 days ago
  •  ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the... 
    Senior

    Webex Events (formerly Socio)

    San Jose, CA
    1 day ago
  • $104.9k - $174.7k

     ...Site Reliability Engineer The Site Reliability Engineer role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and... 
    Senior
    Temporary work
    Local area

    Lexus Nexus

    San Jose, CA
    24 minutes ago
  • $150.4k - $277.6k

     ...Senior Site Reliability Engineer The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple... 
    Senior
    Relocation
    Day shift

    Apple

    Cupertino, CA
    4 days ago
  •  ...Job Title: Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 10+ years of experience in SRE, DevOps, or infrastructure engineering Strong experience... 
    Senior
    Full time

    SFE

    San Jose, CA
    2 days ago
  •  ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 8+ years of experience in SRE, DevOps, or platform engineering Hands-on experience... 
    Senior
    Full time

    SFE

    San Jose, CA
    2 days ago
  • $165k - $265k

     ...possibilities of AI. Role Overviewd-Matrix is looking for a Senior Staff Power Performance Architect to own pre-silicon power estimation...  .... You will partner closely with front-end architects, DV engineers, and backend design teams to identify power activity windows, define... 
    Senior

    d-Matrix

    Santa Clara, CA
    a month ago
  • $248k - $396.75k

     ...US, CA, Santa Clara Full time JR2023973 Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $151k - $244.2k

     ..., you will collaborate closely with our engineering teams to develop innovative solutions that...  ...performance and health. As a Senior Staff SRE with the Cortex Observability team,...  ...operability of the product and ensure the reliability and availability of our services.... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  •  ...competitive benchmarking into clear product and portfolio strategy. The role partners with and influences cross-functional teams across engineering, regional marketing, sales, FAEs, and operations to drive execution and business growth. It also contributes to benchmarking,... 
    Senior
    Full time
    Temporary work
    Worldwide

    Renesas Electronics

    San Jose, CA
    25 days ago
  • $158k - $225k

     ...Description Job Description Antora Energy delivers affordable, reliable energy to industry, data centers, and the grid. Our thermal...  ...costs changes, etc. Collaborate cross-functionally with engineering, project finance, and legal teams to gather insights and resources... 
    Senior
    Flexible hours

    Antora Energy

    San Jose, CA
    26 days ago
  •  ...runs on complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems remain...  ...engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on... 
    Senior

    Saviynt

    Milpitas, CA
    a month ago
  •  ...Job Description Job Description Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots... 
    Permanent employment
    Full time
    Work at office
    Local area

    Foxconn Industrial Internet - FII

    San Jose, CA
    a month ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    a month ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    a month ago
  •  ...Must Have Technical/Functional Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments. Practical vulnerability-management experience; familiarity with... 
    Full time
    Worldwide

    SFE

    San Jose, CA
    3 days ago
  •  ...The RoleThis hybrid role combines the hands-on responsibilities of a Technical Support Engineer within a SaaS (Software as a Service) environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate has a strong technical foundation, thrives in a... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    15 days ago
  • $190.9k - $334.1k

     ...It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful... 
    Senior
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    20 days ago
  • $60 - $62 per hour

     ...and improving existing processes to enhance overall system reliability. Key Responsibilities: Deploying software to cloud...  ...Computer Science or a related field. 3+ years of experience in Site Reliability Engineering. Proficiency with Kubernetes, Helm, Linux, AWS networking... 
    Hourly pay
    Contract work
    Remote work

    Akraya

    Santa Clara, CA
    2 days ago
  • $110k - $175k

     ...wide range of distinguished customers globally. Description We are looking for a driven, team-oriented Senior Signal Integrity Engineer to develop next generation Solid State Drive (SSD) products. This position requires experience in hardware design and signal and... 
    Senior
    Work experience placement
    Flexible hours

    SK hynix memory solutions America Inc.

    Santa Clara, CA
    18 days ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support our client's team based out of... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    3 days ago
  • $110k - $130k

     ...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global...  ...every day? Join us at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    1 day ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    San Jose, CA
    1 day ago
  • $122.5k - $175k

     ...impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San Jose, CA office 3 days a week, reporting to the Chief... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    2 days ago
  • $180k - $225k

     ...Zscaler.RoleWe are looking for a Senior Staff Rust Developer to join our Platform...  ...in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from...  ...of millions of users with high reliability and low latency. You will design and... 
    Senior
    Full time
    Work at office
    Local area

    Zscaler

    San Jose, CA
    21 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Staff Site Reliability Engineer. Be the first to apply!