Site Reliability Engineer
RecruiterPerry
This role requires candidates who are currently authorized to work in the U.S. without sponsorship, and C2C arrangements are not accepted. This role is onsite near Tustin, CA. Job Description
We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions. The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability , along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams. Responsibilities
We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions. The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability , along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams. Responsibilities
- Build, manage, maintain, and troubleshoot Kubernetes clusters and containerized application environments.
- Partner closely with product and application development teams to understand infrastructure requirements and translate them into actionable platform engineering needs.
- Serve as a first point of contact for infrastructure and reliability issues affecting product teams, performing root-cause analysis across application, Kubernetes, cloud, networking, and security layers.
- Resolve Kubernetes and platform-related issues directly while coordinating with other infrastructure, cloud, network, and security teams when issues fall outside the platform team's ownership.
- Create, configure, and maintain AWS resources supporting application and platform environments.
- Support and enhance an internal observability platform used to monitor applications, services, and infrastructure.
- Onboard new applications and use cases into the observability platform by partnering with technical teams to understand monitoring and telemetry requirements.
- Develop and maintain Python scripts used for infrastructure automation, troubleshooting, platform operations, and observability.
- Support CI/CD and GitOps-based deployment processes for Kubernetes environments.
- Improve platform reliability, scalability, monitoring, operational efficiency, and developer experience.
- Participate in troubleshooting and root-cause investigations involving multiple engineering teams and drive issues through successful resolution.
- Document platform standards, troubleshooting procedures, operational processes, and technical solutions.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience , or 6+ years of equivalent professional experience in lieu of a degree.
- Deep hands-on experience with Kubernetes , including building clusters from scratch, cluster administration, deployments, networking, and troubleshooting.
- Strong hands-on experience with AWS , including creating and maintaining cloud infrastructure and resources.
- Experience working with AWS services such as S3 and RDS .
- Strong understanding of observability, monitoring, logging, metrics, and application performance concepts .
- Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus , or similar technologies.
- Experience using Python for scripting, automation, or infrastructure-related tasks.
- Strong troubleshooting and root-cause analysis skills across complex application and infrastructure environments.
- Excellent verbal and written communication skills with the ability to work effectively across product, application, infrastructure, cloud, networking, and security teams.
- Ability and willingness to work in a highly collaborative position that combines hands-on engineering with significant cross-team coordination.
- Experience with Helm for Kubernetes application packaging and deployment.
- Experience with Argo CD and GitOps-based deployment practices.
- Experience building or maintaining CI/CD pipelines .
- Experience with Prometheus, Grafana, and OpenTelemetry .
- Experience with Amazon EKS and IAM .
- Familiarity with Istio or other service mesh technologies.
- Experience with Kafka or Amazon MSK .
- Experience with infrastructure-as-code and configuration-management technologies such as Terraform or Ansible .
- Experience supporting internal developer platforms, infrastructure platforms, or observability products.
- Experience onboarding development teams or applications onto centralized platform services.
- Experience working in regulated or highly controlled technology environments.
- Interest in learning and taking ownership of application code supporting internal platform tooling.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Tustin, CA vacancy
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Healthcare Payments team, you will solve complex and broad...Suggested
$180k - $230k
...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...SuggestedWork at officeLocal areaImmediate startRemote work3 days per week- ...Job Description Job Description Position Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) with hands-on experience in warehouse automation, industrial control systems, and system integration. This role focuses on ensuring reliability...SuggestedRemote work
- ...Job Description One of Insight Global’s customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support...SuggestedContract work
- ...Taco Bell Digital Site Reliability Engineer Taco Bell was born and raised in California and has been around since 1962. We went from selling everyone's favorite Crunchy Tacos on the West Coast to a global brand with 7,500+ restaurants, 350 franchise organizations, that...Work experience placementWork at officeRemote workFlexible hours
$100k - $140k
...and software services. Our team of passionate engineers are constantly innovating, engineering... ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced Site Reliability Engineer to join our team and play...Local areaWorldwideWeekend work$105k - $127k
...Third-party recruiters: Please do not contact us about this role. Site Reliability Engineer At ICEYE, we design, build and operate the largest fleet of Synthetic Aperture Radar (SAR) satellites in the world. Using advanced technology, our constellation collects topographical...Full timeWork at officeLocal areaFlexible hoursNight shift$192.4k - $275.8k
...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...Full timeTemporary workLocal areaFlexible hours$113.3k - $205.52k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...Work at officeRemote workWorldwideFlexible hours$90k - $130k
...discipline. ~5–7 years of practical experience in warehouse automation, industrial control systems, or a similar Site Reliability/Systems Engineering position. ~ Strong command of PLC programming, SCADA/MES environments, and robotics integration. ~ Proficiency in...Full timeRemote work$55k - $187k
...Internal Firm Services - Other Management Level Senior Associate Job Description & Summary The Opportunity As a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$102.8k - $190.2k
...Senior Site Reliability Engineer, Data & Analytics This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large-scale data...Full timeTemporary workPart timeLocal areaRemote workRelocation package$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...Full timeWork experience placementImmediate start$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa...Full timeWork experience placementImmediate start$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start$145.7k - $218.5k
...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to... ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated...Work experience placementShift work$166k - $250k
...to serve a wide variety of defense, IC and commercial customers in US and international markets.ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems...Full timeWork experience placementImmediate start$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine...Full timeWork experience placementImmediate start$191k - $253k
...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate... ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril... ...:10+ years of experience in site reliability engineering, production engineering...Full timeWork experience placementImmediate start$253k - $336k
...not years.ABOUT THE TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril's corporate systems.It drives... ...its mission critical products.ABOUT THE JOB:The Director of Site Reliability Engineering owns the reliability system for the software that...Full timeWork experience placementImmediate start$37.26 - $68.93 per hour
Team Name:DiabloJob Title:Software Engineer, Engine SystemsRequisition ID:R027570Job Description... ...work week, with a mix of remote and on-site days. While hybrid is the standard... ..., fixing bugs, and strengthening system reliability. Advance the game engine through performance...Hourly payFull timeTemporary workPart timeLocal areaRemote workRelocation packageFlexible hours$150k - $180k
WHAT YOU’LL DOThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess a deep understanding of systems engineering and automation, including configuration management...Work experience placementLocal area$103.17k - $158.87k
10883 - Enterprise Systems & AI Solutions Engineer Irvine, CA Company Overview Hyundai AutoEver America (HAEA) is the dynamic... ...IT strategy. Maintain enterprise data integrity, system reliability, integrations, and user support processes. Design and develop...Local area$167.67k - $204.93k
...a product that meets the high level of safety and reliability in today’s air transportation system.The future of... ...core values. This position is required to work on-site 5 days a week.What we do: As our Lead Engineer, System Safety, you will take ownership of safety engineering...Local area- ...and the guarantee that every piece of hardware on the board comes up and works when our image boots. We’re hiring a Senior Software Engineer to own that layer.This role sits at the seam between our platform organization, which maintains the underlying NixOS platform and...Full timeWork experience placementLive inImmediate start
$88.97k - $244.51k
...Retail & Consumer Goods, Communications Media & Technology, Consumer Business Services, Public Sector and others..OverviewThe Solution Engineer is a role that we often hire at Salesforce. If you are interested in any type of Solution Engineering role, you've come to the...Full timeWork at officeRemote workFlexible hours$202.4k - $257.9k
...are seeking a proactive, highly technical, and AI-first Solutions Engineer to step in and co-own our business in the region.Your ImpactAs a... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to...Full timeTemporary workLocal areaRelocationFlexible hoursShift work- Job Summary:The Engineer II, Cloud & Web Systems Software is responsible for the design, development... ...the development of scalable, secure, and reliable applications and services that enable... ...Assistance Program, Pet Insurance, on-site Wellness Clinic, Fitness Center, Café....Work at officeFlexible hours
$63k - $140k
...ApplicableSpecialismData, Analytics & AIManagement LevelAssociateJob Description & SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage advanced technologies and techniques to design and develop robust data solutions for...Full timeH1b
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- IT site lead Tustin, CA
- site leader Tustin, CA
- on-site clinical research associate (traveling/remote) Tustin, CA
- construction site safety Tustin, CA
- site services specialist Tustin, CA
- site reliability engineer sre
- site reliability engineer
- site reliability engineering manager
- site reliability engineer remote
- junior site reliability engineer


