Site Reliability Engineer
RecruiterPerry
This role requires candidates who are currently authorized to work in the U.S. without sponsorship, and C2C arrangements are not accepted. This role is onsite near Tustin, CA. Job Description
We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions. The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability , along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams. Responsibilities
We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions. The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability , along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams. Responsibilities
- Build, manage, maintain, and troubleshoot Kubernetes clusters and containerized application environments.
- Partner closely with product and application development teams to understand infrastructure requirements and translate them into actionable platform engineering needs.
- Serve as a first point of contact for infrastructure and reliability issues affecting product teams, performing root-cause analysis across application, Kubernetes, cloud, networking, and security layers.
- Resolve Kubernetes and platform-related issues directly while coordinating with other infrastructure, cloud, network, and security teams when issues fall outside the platform team's ownership.
- Create, configure, and maintain AWS resources supporting application and platform environments.
- Support and enhance an internal observability platform used to monitor applications, services, and infrastructure.
- Onboard new applications and use cases into the observability platform by partnering with technical teams to understand monitoring and telemetry requirements.
- Develop and maintain Python scripts used for infrastructure automation, troubleshooting, platform operations, and observability.
- Support CI/CD and GitOps-based deployment processes for Kubernetes environments.
- Improve platform reliability, scalability, monitoring, operational efficiency, and developer experience.
- Participate in troubleshooting and root-cause investigations involving multiple engineering teams and drive issues through successful resolution.
- Document platform standards, troubleshooting procedures, operational processes, and technical solutions.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience , or 6+ years of equivalent professional experience in lieu of a degree.
- Deep hands-on experience with Kubernetes , including building clusters from scratch, cluster administration, deployments, networking, and troubleshooting.
- Strong hands-on experience with AWS , including creating and maintaining cloud infrastructure and resources.
- Experience working with AWS services such as S3 and RDS .
- Strong understanding of observability, monitoring, logging, metrics, and application performance concepts .
- Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus , or similar technologies.
- Experience using Python for scripting, automation, or infrastructure-related tasks.
- Strong troubleshooting and root-cause analysis skills across complex application and infrastructure environments.
- Excellent verbal and written communication skills with the ability to work effectively across product, application, infrastructure, cloud, networking, and security teams.
- Ability and willingness to work in a highly collaborative position that combines hands-on engineering with significant cross-team coordination.
- Experience with Helm for Kubernetes application packaging and deployment.
- Experience with Argo CD and GitOps-based deployment practices.
- Experience building or maintaining CI/CD pipelines .
- Experience with Prometheus, Grafana, and OpenTelemetry .
- Experience with Amazon EKS and IAM .
- Familiarity with Istio or other service mesh technologies.
- Experience with Kafka or Amazon MSK .
- Experience with infrastructure-as-code and configuration-management technologies such as Terraform or Ansible .
- Experience supporting internal developer platforms, infrastructure platforms, or observability products.
- Experience onboarding development teams or applications onto centralized platform services.
- Experience working in regulated or highly controlled technology environments.
- Interest in learning and taking ownership of application code supporting internal platform tooling.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Tustin, CA vacancy
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Healthcare Payments team, you will solve complex and broad...Suggested
- ...One of Insight Global's customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support deployments,...SuggestedContract work
$100k - $140k
...and software services. Our team of passionate engineers are constantly innovating, engineering... ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced Site Reliability Engineer to join our team and play...SuggestedLocal areaWorldwideWeekend work- ...Role:- Site Reliability Engineer (SRE) Location:- Irvine, CA Duration:- 12 Months Role Description: 1) Develop and provide operational support for full-stack software applications. 2) Collaborate with development operations staff to create, monitor, and troubleshoot the...Suggested
$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours$102.8k - $190.2k
...Team Name: Battle.net & Online Products Job Title: Senior Site Reliability Engineer, Data & Analytics Requisition ID: R027436 Job Description: This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics...Full timeTemporary workPart timeLocal areaRemote workRelocation package$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa...Full timeWork experience placementImmediate start$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...Full timeWork experience placementImmediate start$166k - $250k
...to serve a wide variety of defense, IC and commercial customers in US and international markets.ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems...Full timeWork experience placementImmediate start$145.7k - $218.5k
...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to... ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated...Work experience placementShift work$166k - $220k
...technology to the military in months, not years. ABOUT THE TEAM We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability,...Full timeWork experience placementImmediate startRemote work$191k - $253k
...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate... ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril... ...:10+ years of experience in site reliability engineering, production engineering...Full timeWork experience placementImmediate start$253k - $336k
...not years.ABOUT THE TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril's corporate systems.It drives... ...its mission critical products.ABOUT THE JOB:The Director of Site Reliability Engineering owns the reliability system for the software that...Full timeWork experience placementImmediate start$158.6k - $237.6k
...End strategy and PD execution for a high-quality SS design delivery to SOC customers.What You Can ExpectThe IP Release Principal Engineer is a senior technical individual contributor within Marvell's IP Subsystem (IPSS) Center of Excellence (COE). This role is the quality...Permanent employmentInternshipWork from home$37.26 - $68.93 per hour
Team Name:DiabloJob Title:Software Engineer, Engine SystemsRequisition ID:R027570Job Description... ...work week, with a mix of remote and on-site days. While hybrid is the standard... ..., fixing bugs, and strengthening system reliability. Advance the game engine through performance...Hourly payFull timeTemporary workPart timeLocal areaRemote workRelocation packageFlexible hours$150k - $180k
WHAT YOU’LL DOThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess a deep understanding of systems engineering and automation, including configuration management...Work experience placementLocal area$107.8k - $134.8k
...Senior Mechanical Engineer – Exterior TrimRivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we constantly challenge what...Permanent employmentFull timeTemporary workPart timeLocal areaShift work$107.8k - $134.8k
...but our team shares a love of the outdoors and a desire to protect it for future generations. Role Summary As a Senior Mechanical Engineer - Exterior Trim you will own the design, development, and release of exterior trim systems and components (e.g., fascias, flares,...Permanent employmentFull timeContract workTemporary workPart timeLocal areaShift work$102k - $173.4k
...Team: Advancing solution enablement, business practice maturity, and partner readiness. (IBM Software)As a Sr Technical Solutions Engineer, you need to be a technical leader and resource for SD&S and Ingram Micro. Maintain relevant applicable industry certifications and...Full timeTemporary workWork at officeRemote workWorldwideShift work$167.67k - $204.93k
...a product that meets the high level of safety and reliability in today’s air transportation system.The future of... ...core values. This position is required to work on-site 5 days a week.What we do: As our Lead Engineer, System Safety, you will take ownership of safety engineering...Local area$46.9 - $75.04 per hour
Splunk Administrator (Site Reliability Engineer) Pay Range: $46.90 - $75.04 per hour Scheduled Weekly Hours: 40 Responsibilities Deploy, manage, and optimize the Splunk platform for enterprise applications, security operations, and observability practices. Onboard machine...Hourly payRemote work- ...and the guarantee that every piece of hardware on the board comes up and works when our image boots. We’re hiring a Senior Software Engineer to own that layer.This role sits at the seam between our platform organization, which maintains the underlying NixOS platform and...Full timeWork experience placementLive inImmediate start
$103.17k - $158.87k
10883 - Enterprise Systems & AI Solutions Engineer Irvine, CA Company Overview Hyundai AutoEver America (HAEA) is the dynamic... ...IT strategy. Maintain enterprise data integrity, system reliability, integrations, and user support processes. Design and develop...Local area$88.97k - $244.51k
...Retail & Consumer Goods, Communications Media & Technology, Consumer Business Services, Public Sector and others..OverviewThe Solution Engineer is a role that we often hire at Salesforce. If you are interested in any type of Solution Engineering role, you've come to the...Full timeWork at officeRemote workFlexible hours- Job Summary:The Engineer II, Cloud & Web Systems Software is responsible for the design, development... ...the development of scalable, secure, and reliable applications and services that enable... ...Assistance Program, Pet Insurance, on-site Wellness Clinic, Fitness Center, Café....Work at officeFlexible hours
$63k - $140k
...ApplicableSpecialismData, Analytics & AIManagement LevelAssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage advanced technologies and techniques to design and develop robust data solutions for...Full timeH1b$191k - $253k
...customers to rapidly close the kill chain against a broad range of Unmanned Aerial System (UAS) threats. Working across product, engineering, sales, logistics, operations, and mission success, the Air Defense team develops, tests, deploys, and sustains the Anduril Air Defense...Full timeWork experience placementImmediate startRemote workWorldwide$220k - $292k
...common operating picture. About the Job We're looking for Software Engineer specializing in Robotics to join our growing team in Irvine,... ...Be responsible for service ownership, ensuring functionality, reliability, and alignment with customer objectives. This includes...Full timeWork experience placementLive inImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- on-site clinical research associate (traveling/remote) Tustin, CA
- construction site safety Tustin, CA
- junior site reliability engineer
- site reliability engineer remote
- lead site reliability engineer
- site reliability engineer
- site reliability engineer sre
- site reliability engineering manager
- junior website developer
- website auditor

