Site Reliability Engineer
Castelion
Why Castelion, Why Now
Castelion is moving incredibly fast to develop and deliver advanced defense systems at a time when execution matters more than ever. We believe focus, ownership, and excellence are decisive advantages - and we're building a world-class team to turn bold ideas into real capability.
This is a rare opportunity to join at an early stage, where you'll have significant ownership, collaborate with exceptional teammates, and make a direct, measurable impact on our mission and the future of the company - regardless of your function.
Site Reliability Engineer
We are seeking a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.
This role is the missing reliability piece of an existing high-performing engineering organization. You will work across DevOps, Cloud, Software, Security, Test, and IT to identify reliability risks, diagnose failures that cross system boundaries, and drive corrective actions to resolution. You will be expected to understand and improve existing systems rather than defaulting to replacement, using new technology when it solves a demonstrated reliability, scalability, or operational problem.
Responsibilities
- Establish meaningful reliability, availability, latency, capacity, and recovery expectations for critical engineering services, with measurable health indicators and useful alerts.
- Lead deep technical investigations and incident response across application, Linux, networking, storage, Kubernetes, cloud, and other system boundaries; collect evidence, separate symptoms from root causes, and drive incidents through resolution.
- Build and improve monitoring and diagnostic systems that detect problems before users report them and provide engineers with the information needed to quickly understand and resolve failures.
- Analyze system performance and capacity across compute, memory, storage, networking, connections, and other constrained resources; identify operating limits and address issues through the simplest effective solution, whether optimization, additional capacity, scaling, caching, configuration changes, or architectural improvements.
- Drive evidence-backed root cause analysis and postmortem actions for significant incidents, ensuring corrective and preventive actions are implemented and verified to reduce recurring failures.
- Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve reliability problems that cross team boundaries, providing technical leadership without attempting to own every component involved.
- Understand, operate, and incrementally improve systems built by other engineers, balancing reliability and operational value against existing architecture, constraints, and engineering practices.
- Participate in the on-call rotation for critical engineering services, providing first-response triage, escalation, and follow-up for recurring reliability issues.
Basic Qualifications
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
- Demonstrated experience debugging complex production failures across multiple system layers and driving investigations from the first symptom to an evidence-backed root cause and lasting corrective action.
- Strong Linux systems expertise, including CPU, memory, storage, networking, processes, sockets, and system services, with a strong understanding of performance and capacity concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
- Experience building and operating observability, monitoring, alerting, and incident response systems, with the ability to distinguish between mitigation, workaround, corrective action, and preventive action.
- Strong networking and application fundamentals, including TCP, TLS, DNS, reverse proxies, load balancers, connection states, and timeouts; able to investigate application runtime behavior such as threads, connection pools, file descriptors, memory, or garbage collection when the evidence points there.
- Demonstrated ability to work effectively within existing systems and across engineering organizations, asking why a system was designed a certain way and improving it based on measurable reliability and operational needs rather than defaulting to rewrites or replacement.
Castelion offers a generous benefits package. Please refer to the bottom of our Careers page for more details.
Other Duties
Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.
Additional Eligibility Requirements
This position may require access to classified information or restricted U.S. Government sites, systems, or information, as determined by the Company and/or applicable U.S. Government requirements. If the position is so designated, your employment in the role may be contingent upon your ability to obtain and maintain the required U.S. Government security clearance or other government authorization, and to satisfy any citizenship or other eligibility requirements imposed by applicable law, regulation, executive order, or government contract requirements. You will be notified if and when such requirements apply.
Affirmative Action/EEO Statement
Castelion is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.
Castelion is committed to providing reasonable accommodations to qualified individuals with disabilities throughout the application and hiring process. If you require a reasonable accommodation to complete an application, participate in the interview process, or otherwise participate in the hiring process, please contact View email address on click.appcast.io. Requests for accommodation will be considered on an individual basis and handled in accordance with applicable law.
Castelion is committed to fostering a workplace where employment decisions are based on qualifications, business needs, and the ability to perform the essential functions of the role, with or without reasonable accommodation.
EAR/ITAR Requirements
This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.
#J-18808-Ljbffr$155k - $195k
...you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleWe are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site...SuggestedPermanent employmentFull timeWork at office$81.5k - $141.3k
...generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Title: Site Reliability Engineer II Team: CRM DevOps Employment Type: Full-Time About the Role This Site Reliability Engineer II position works...SuggestedFull timeRemote workShift work$180k - $200k
...to you through an Ateme solution created by our award-winning engineering teams. Ateme (PARIS: ATEME) is the global leader in video... ...Culture: Collaborate with talented international teams that value reliability, innovation, knowledge sharing, and continuous improvement....Suggested- ...join us on our mission of providing humankind access to the galaxy beyond our planet. About the Role We are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site...SuggestedFull timeWork at office
$140k - $180k
...fundamentally different class of spacecraft. Engineered to survive the harshest radiation... ...create highly available, deployable, and reliable products Reduce operational toil through... ...experience in Software Engineering, Site Reliability Engineering or DevOps ~ Deep...SuggestedPermanent employmentShift work- ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus...Full timeLocal areaWorldwideFlexible hours
$90k - $180k
...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division. We...Remote workShift work$160k - $200k
...what it means for businesses and their employees to truly feel safe. POSITION OVERVIEW: HiveWatch is seeking a Senior Site Reliability Engineer to join our Platform Team, where you'll build and operate mission-critical edge infrastructure that connects our SaaS...Flexible hours$107.8k - $162k
...an expectation of a minimum of three days per week working in the office and flexibility to work remotely on the remaining days. On-site expectations may evolve over time to support business needs, with clear communication provided in advance. Job Description Operates...Work at officeLocal areaRemote work3 days per week$164k - $270k
...exactly who we’re looking for.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to... ...automated remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop...Permanent employmentFull timeRelocation packageFlexible hours$165k - $265k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation...Permanent employmentTemporary workImmediate startWeekend work$175k - $285k
Hadrian - Manufacturing the FutureHadrian is building autonomous factories to reindustrialize America. By combining AI, advanced software, robotics, and full-stack manufacturing, we help aerospace and defense companies build rockets, satellites, aircraft, ships, and other...Permanent employmentFull timeRemote workRelocation packageFlexible hours$181k - $265k
...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to... ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree...Work at officeImmediate start3 days per week$210.5k - $263.1k
...anime content we all love. Join our team, and help us shape the future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability,...Flexible hours- ...Pivotal Health, a healthcare technology platform, is seeking a Staff Site Reliability Engineer to embed reliability and operational excellence into our production systems. You will lead hands-on engineering across cloud, observability, incident response, and automation...
$100k - $200k
...backed by top tier investors. Our lean, world-class team of engineers and operators is applying a first-principles approach to... ...culture of urgency, accountability and transparency. DevOps / Site Reliability Engineer We are seeking a highly capable DevOps / Site...Full timeWeekend work$230k - $260k
...operate with clarity, control, and confidence across the reimbursement journey. About The Role We’re hiring a Staff Site Reliability Engineer to define and strengthen how reliability, scalability, and operational excellence are built into Pivotal’s platform. This...Remote workFlexible hours- ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating... ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms...Full timeContract workImmediate startWork from homeFlexible hours
$159.8k - $235k
...infrastructure troubleshooting.About the RoleAs a Senior Software Engineer on Code Scalability, you will solve the engineering problems... ...large monorepos.Improve Buildkite pipeline throughput and reliability by reducing queue times, addressing build bottlenecks, and making...Hourly payWork at officeLocal areaRemote workFlexible hours- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$70 - $99 per hour
...Job Description We have an immediate need for a remote FMS/Avionics Software Engineer. This is a contract position for an immediate program. Pay Rate: $70-$99/hr. DOE Job Requirements Due to compliance with U.S. export control laws and regulations...Permanent employmentContract workImmediate startRemote work$208k - $263k
...designs into repeatable, serialized production builds without losing engineering intent between revisions. Keep design iteration fast while... ...build states move through the fab floor, in inventory, and at site simultaneously. Integrate power, cooling, and structural...$135k - $175k
...Database Reliability EngineerHawthorne, CASpaceX was founded under the belief that a future where humanity is out exploring the stars is... ...RELIABILITY ENGINEERSpaceX is looking for a Database Reliability Engineer with strong technical knowledge in Microsoft SQL/PowerBI...Permanent employmentTemporary workRemote workFlexible hoursWeekend work$127k - $184k
...architecture in a customer-facing or support role.Experience with cloud engineering, on-premise engineering, virtualization, or containerization... ...built in the cloud. Our products are developed for security, reliability and scalability, running the full stack from infrastructure...$135.4k - $181.6k
...at the heart of Disney's past, present, and future. Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological...- Riot engineers bring deep knowledge of specific technical areas, but also value the opportunity to work in a variety of broader domains... ..., designers, and other cross-functional teammates to build reliable, scalable, and maintainable Client capabilities. You will help...Local areaFlexible hours
$165k - $230k
...the ultimate goal of enabling human life on Mars.SR. SOFTWARE ENGINEER, PLATFORMThe application software team is the central nervous system... ...as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. We are looking for engineers who treat fellow...Permanent employmentTemporary workWorldwideWeekend work$130.6k - $192k
About the TeamThe Experimentation Platform Team develops a state-of-the-art platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-technical audiences to come up with hypotheses; design, configure and analyze experiments; and conduct...Hourly payWork at officeLocal areaRemote workFlexible hours- ...Lead Systems EngineerDisney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological backbone for Disney's media business...
$138k - $167k
...tools Contextual knowledge of AI technologies and it’s practical implementations in real world and how it applies to system engineering and operations Project management: Waterfall/Agile, ROI, FRD/BRD/TRD, CBA/budgeting Experience working with RDMS, No-SQL...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Los Angeles, CA
- site safety Los Angeles, CA
- website coordinator Los Angeles, CA
- on-site clinical research associate (traveling/remote) Los Angeles, CA
- site services specialist Los Angeles, CA
- on site coordinator Los Angeles, CA
- construction site safety Los Angeles, CA
- junior website developer Los Angeles, CA
- site recruiter Los Angeles, CA
- historic site Los Angeles, CA




