Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Castelion

Why Castelion, Why Now

Castelion is moving incredibly fast to develop and deliver advanced defense systems at a time when execution matters more than ever. We believe focus, ownership, and excellence are decisive advantages - and we're building a world-class team to turn bold ideas into real capability.

This is a rare opportunity to join at an early stage, where you'll have significant ownership, collaborate with exceptional teammates, and make a direct, measurable impact on our mission and the future of the company - regardless of your function.

Site Reliability Engineer

We are seeking a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.

This role is the missing reliability piece of an existing high-performing engineering organization. You will work across DevOps, Cloud, Software, Security, Test, and IT to identify reliability risks, diagnose failures that cross system boundaries, and drive corrective actions to resolution. You will be expected to understand and improve existing systems rather than defaulting to replacement, using new technology when it solves a demonstrated reliability, scalability, or operational problem.

Responsibilities
  • Establish meaningful reliability, availability, latency, capacity, and recovery expectations for critical engineering services, with measurable health indicators and useful alerts.
  • Lead deep technical investigations and incident response across application, Linux, networking, storage, Kubernetes, cloud, and other system boundaries; collect evidence, separate symptoms from root causes, and drive incidents through resolution.
  • Build and improve monitoring and diagnostic systems that detect problems before users report them and provide engineers with the information needed to quickly understand and resolve failures.
  • Analyze system performance and capacity across compute, memory, storage, networking, connections, and other constrained resources; identify operating limits and address issues through the simplest effective solution, whether optimization, additional capacity, scaling, caching, configuration changes, or architectural improvements.
  • Drive evidence-backed root cause analysis and postmortem actions for significant incidents, ensuring corrective and preventive actions are implemented and verified to reduce recurring failures.
  • Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve reliability problems that cross team boundaries, providing technical leadership without attempting to own every component involved.
  • Understand, operate, and incrementally improve systems built by other engineers, balancing reliability and operational value against existing architecture, constraints, and engineering practices.
  • Participate in the on-call rotation for critical engineering services, providing first-response triage, escalation, and follow-up for recurring reliability issues.
Basic Qualifications
  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
  • Demonstrated experience debugging complex production failures across multiple system layers and driving investigations from the first symptom to an evidence-backed root cause and lasting corrective action.
  • Strong Linux systems expertise, including CPU, memory, storage, networking, processes, sockets, and system services, with a strong understanding of performance and capacity concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
  • Experience building and operating observability, monitoring, alerting, and incident response systems, with the ability to distinguish between mitigation, workaround, corrective action, and preventive action.
  • Strong networking and application fundamentals, including TCP, TLS, DNS, reverse proxies, load balancers, connection states, and timeouts; able to investigate application runtime behavior such as threads, connection pools, file descriptors, memory, or garbage collection when the evidence points there.
  • Demonstrated ability to work effectively within existing systems and across engineering organizations, asking why a system was designed a certain way and improving it based on measurable reliability and operational needs rather than defaulting to rewrites or replacement.

Castelion offers a generous benefits package. Please refer to the bottom of our Careers page for more details.

Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.

Additional Eligibility Requirements

This position may require access to classified information or restricted U.S. Government sites, systems, or information, as determined by the Company and/or applicable U.S. Government requirements. If the position is so designated, your employment in the role may be contingent upon your ability to obtain and maintain the required U.S. Government security clearance or other government authorization, and to satisfy any citizenship or other eligibility requirements imposed by applicable law, regulation, executive order, or government contract requirements. You will be notified if and when such requirements apply.

Affirmative Action/EEO Statement

Castelion is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.

Castelion is committed to providing reasonable accommodations to qualified individuals with disabilities throughout the application and hiring process. If you require a reasonable accommodation to complete an application, participate in the interview process, or otherwise participate in the hiring process, please contact View email address on click.appcast.io. Requests for accommodation will be considered on an individual basis and handled in accordance with applicable law.

Castelion is committed to fostering a workplace where employment decisions are based on qualifications, business needs, and the ability to perform the essential functions of the role, with or without reasonable accommodation.

EAR/ITAR Requirements

This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.

#J-18808-Ljbffr
Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Los Angeles, CA vacancy
  • $155k - $195k

     ...you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleWe are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site... 
    Suggested
    Permanent employment
    Full time
    Work at office

    Apex Technology

    Los Angeles, CA
    3 days ago
  • $81.5k - $141.3k

     ...generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Title: Site Reliability Engineer II Team: CRM DevOps Employment Type: Full-Time About the Role This Site Reliability Engineer II position works... 
    Suggested
    Full time
    Remote work
    Shift work

    Talentify.io

    Los Angeles, CA
    6 hours ago
  • $180k - $200k

     ...to you through an Ateme solution created by our award-winning engineering teams. Ateme (PARIS: ATEME) is the global leader in video...  ...Culture: Collaborate with talented international teams that value reliability, innovation, knowledge sharing, and continuous improvement.... 
    Suggested

    ATEME

    Culver City, CA
    3 days ago
  •  ...join us on our mission of providing humankind access to the galaxy beyond our planet. About the Role We are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site... 
    Suggested
    Full time
    Work at office

    Apex Space, Inc.

    Los Angeles, CA
    3 days ago
  • $140k - $180k

     ...fundamentally different class of spacecraft. Engineered to survive the harshest radiation...  ...create highly available, deployable, and reliable products Reduce operational toil through...  ...experience in Software Engineering, Site Reliability Engineering or DevOps ~ Deep... 
    Suggested
    Permanent employment
    Shift work

    K2 Space

    Los Angeles, CA
    3 days ago
  •  ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus... 
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group Inc

    Los Angeles, CA
    5 days ago
  • $90k - $180k

     ...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division. We... 
    Remote work
    Shift work

    Talentify.io

    Los Angeles, CA
    6 hours ago
  • $160k - $200k

     ...what it means for businesses and their employees to truly feel safe. POSITION OVERVIEW: HiveWatch is seeking a Senior Site Reliability Engineer to join our Platform Team, where you'll build and operate mission-critical edge infrastructure that connects our SaaS... 
    Flexible hours

    HiveWatch

    Los Angeles, CA
    3 days ago
  • $107.8k - $162k

     ...an expectation of a minimum of three days per week working in the office and flexibility to work remotely on the remaining days. On-site expectations may evolve over time to support business needs, with clear communication provided in advance. Job Description Operates... 
    Work at office
    Local area
    Remote work
    3 days per week

    Green Dot

    Los Angeles, CA
    1 day ago
  • $164k - $270k

     ...exactly who we’re looking for.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to...  ...automated remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop... 
    Permanent employment
    Full time
    Relocation package
    Flexible hours

    Hadrian

    Los Angeles, CA
    6 hours ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $175k - $285k

    Hadrian - Manufacturing the FutureHadrian is building autonomous factories to reindustrialize America. By combining AI, advanced software, robotics, and full-stack manufacturing, we help aerospace and defense companies build rockets, satellites, aircraft, ships, and other...
    Permanent employment
    Full time
    Remote work
    Relocation package
    Flexible hours

    Hadrian

    Los Angeles, CA
    1 day ago
  • $181k - $265k

     ...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to...  ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    Los Angeles, CA
    1 day ago
  • $210.5k - $263.1k

     ...anime content we all love. Join our team, and help us shape the future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability,... 
    Flexible hours

    Ellation

    Los Angeles, CA
    2 days ago
  •  ...Pivotal Health, a healthcare technology platform, is seeking a Staff Site Reliability Engineer to embed reliability and operational excellence into our production systems. You will lead hands-on engineering across cloud, observability, incident response, and automation... 

    Pivotal Health

    Santa Monica, CA
    2 days ago
  • $100k - $200k

     ...backed by top tier investors. Our lean, world-class team of engineers and operators is applying a first-principles approach to...  ...culture of urgency, accountability and transparency. DevOps / Site Reliability Engineer We are seeking a highly capable DevOps / Site... 
    Full time
    Weekend work

    General Matter

    Los Angeles, CA
    5 days ago
  • $230k - $260k

     ...operate with clarity, control, and confidence across the reimbursement journey. About The Role We’re hiring a Staff Site Reliability Engineer to define and strengthen how reliability, scalability, and operational excellence are built into Pivotal’s platform. This... 
    Remote work
    Flexible hours

    Pivotal Health

    Los Angeles, CA
    6 hours ago
  •  ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating...  ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms... 
    Full time
    Contract work
    Immediate start
    Work from home
    Flexible hours

    KēSTA I.T.

    Beverly Hills, CA
    18 days ago
  • $159.8k - $235k

     ...infrastructure troubleshooting.About the RoleAs a Senior Software Engineer on Code Scalability, you will solve the engineering problems...  ...large monorepos.Improve Buildkite pipeline throughput and reliability by reducing queue times, addressing build bottlenecks, and making... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Los Angeles, CA
    2 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Los Angeles, CA
    7 days ago
  • $70 - $99 per hour

     ...Job Description  We have an immediate need for a remote FMS/Avionics Software Engineer. This is a contract position for an immediate program.    Pay Rate: $70-$99/hr. DOE Job Requirements Due to compliance with U.S. export control laws and regulations... 
    Permanent employment
    Contract work
    Immediate start
    Remote work

    Performance Software Corporation

    Gardena, CA
    1 day ago
  • $208k - $263k

     ...designs into repeatable, serialized production builds without losing engineering intent between revisions. Keep design iteration fast while...  ...build states move through the fab floor, in inventory, and at site simultaneously. Integrate power, cooling, and structural... 

    Fluidstack

    Los Angeles, CA
    6 hours ago
  • $135k - $175k

     ...Database Reliability EngineerHawthorne, CASpaceX was founded under the belief that a future where humanity is out exploring the stars is...  ...RELIABILITY ENGINEERSpaceX is looking for a Database Reliability Engineer with strong technical knowledge in Microsoft SQL/PowerBI... 
    Permanent employment
    Temporary work
    Remote work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    5 days ago
  • $127k - $184k

     ...architecture in a customer-facing or support role.Experience with cloud engineering, on-premise engineering, virtualization, or containerization...  ...built in the cloud. Our products are developed for security, reliability and scalability, running the full stack from infrastructure... 

    Google

    Los Angeles, CA
    1 day ago
  • $135.4k - $181.6k

     ...at the heart of Disney's past, present, and future. Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological... 

    Disney

    Glendale, CA
    1 day ago
  • Riot engineers bring deep knowledge of specific technical areas, but also value the opportunity to work in a variety of broader domains...  ..., designers, and other cross-functional teammates to build reliable, scalable, and maintainable Client capabilities. You will help... 
    Local area
    Flexible hours

    Riot Games

    Los Angeles, CA
    4 days ago
  • $165k - $230k

     ...the ultimate goal of enabling human life on Mars.SR. SOFTWARE ENGINEER, PLATFORMThe application software team is the central nervous system...  ...as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. We are looking for engineers who treat fellow... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    23 hours ago
  • $130.6k - $192k

    About the TeamThe Experimentation Platform Team develops a state-of-the-art platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-technical audiences to come up with hypotheses; design, configure and analyze experiments; and conduct... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Los Angeles, CA
    4 days ago
  •  ...Lead Systems EngineerDisney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – all working to build and advance the technological backbone for Disney's media business... 

    The Walt Disney Studios

    Glendale, CA
    1 day ago
  • $138k - $167k

     ...tools Contextual knowledge of AI technologies and it’s practical implementations in real world and how it applies to system engineering and operations Project management: Waterfall/Agile, ROI, FRD/BRD/TRD, CBA/budgeting Experience working with RDMS, No-SQL... 

    Sony Pictures Entertainment

    Culver City, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!