Site Reliability Engineer
Castelion
Why Castelion, Why Now
Castelion is moving incredibly fast to develop and deliver advanced defense systems at a time when execution matters more than ever. We believe focus, ownership, and excellence are decisive advantages - and we're building a world-class team to turn bold ideas into real capability.
This is a rare opportunity to join at an early stage, where you'll have significant ownership, collaborate with exceptional teammates, and make a direct, measurable impact on our mission and the future of the company - regardless of your function.
Site Reliability Engineer
We are seeking a Site Reliability Engineer to own the reliability, performance, observability, and operational health of Castelion's critical engineering systems. These systems support software development, CI/CD, artifact distribution, test infrastructure, developer workflows, and other services that engineers depend on to deliver hardware and software.
This role is the missing reliability piece of an existing high-performing engineering organization. You will work across DevOps, Cloud, Software, Security, Test, and IT to identify reliability risks, diagnose failures that cross system boundaries, and drive corrective actions to resolution. You will be expected to understand and improve existing systems rather than defaulting to replacement, using new technology when it solves a demonstrated reliability, scalability, or operational problem.
Responsibilities
- Establish meaningful reliability, availability, latency, capacity, and recovery expectations for critical engineering services, with measurable health indicators and useful alerts.
- Lead deep technical investigations and incident response across application, Linux, networking, storage, Kubernetes, cloud, and other system boundaries; collect evidence, separate symptoms from root causes, and drive incidents through resolution.
- Build and improve monitoring and diagnostic systems that detect problems before users report them and provide engineers with the information needed to quickly understand and resolve failures.
- Analyze system performance and capacity across compute, memory, storage, networking, connections, and other constrained resources; identify operating limits and address issues through the simplest effective solution, whether optimization, additional capacity, scaling, caching, configuration changes, or architectural improvements.
- Drive evidence-backed root cause analysis and postmortem actions for significant incidents, ensuring corrective and preventive actions are implemented and verified to reduce recurring failures.
- Partner with DevOps, Cloud, Software, Security, Test, and IT to resolve reliability problems that cross team boundaries, providing technical leadership without attempting to own every component involved.
- Understand, operate, and incrementally improve systems built by other engineers, balancing reliability and operational value against existing architecture, constraints, and engineering practices.
- Participate in the on-call rotation for critical engineering services, providing first-response triage, escalation, and follow-up for recurring reliability issues.
Basic Qualifications
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
- Demonstrated experience debugging complex production failures across multiple system layers and driving investigations from the first symptom to an evidence-backed root cause and lasting corrective action.
- Strong Linux systems expertise, including CPU, memory, storage, networking, processes, sockets, and system services, with a strong understanding of performance and capacity concepts such as IOPS, throughput, latency, queue depth, and connection concurrency.
- Experience building and operating observability, monitoring, alerting, and incident response systems, with the ability to distinguish between mitigation, workaround, corrective action, and preventive action.
- Strong networking and application fundamentals, including TCP, TLS, DNS, reverse proxies, load balancers, connection states, and timeouts; able to investigate application runtime behavior such as threads, connection pools, file descriptors, memory, or garbage collection when the evidence points there.
- Demonstrated ability to work effectively within existing systems and across engineering organizations, asking why a system was designed a certain way and improving it based on measurable reliability and operational needs rather than defaulting to rewrites or replacement.
Castelion offers a generous benefits package. Please refer to the bottom of our Careers page for more details.
Other Duties
Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities and activities may change at any time with or without notice.
Additional Eligibility Requirements
This position may require access to classified information or restricted U.S. Government sites, systems, or information, as determined by the Company and/or applicable U.S. Government requirements. If the position is so designated, your employment in the role may be contingent upon your ability to obtain and maintain the required U.S. Government security clearance or other government authorization, and to satisfy any citizenship or other eligibility requirements imposed by applicable law, regulation, executive order, or government contract requirements. You will be notified if and when such requirements apply.
Affirmative Action/EEO Statement
Castelion is an Equal Opportunity Employer. We are committed to providing equal employment opportunities to all applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.
Castelion is committed to providing reasonable accommodations to qualified individuals with disabilities throughout the application and hiring process. If you require a reasonable accommodation to complete an application, participate in the interview process, or otherwise participate in the hiring process, please contact View email address on click.appcast.io. Requests for accommodation will be considered on an individual basis and handled in accordance with applicable law.
Castelion is committed to fostering a workplace where employment decisions are based on qualifications, business needs, and the ability to perform the essential functions of the role, with or without reasonable accommodation.
EAR/ITAR Requirements
This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.
#J-18808-Ljbffr$96.8k - $145.2k
...If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX), United States (US).Job Responsibilities Include: Own and manage...SuggestedTemporary workWork at officeRemote workFlexible hours$65 per hour
...Our client, a leading organization in cloud infrastructure and enterprise solutions, is seeking a Site Reliability Engineer III to join their team. As a Site Reliability Engineer III, you will be part of the Infrastructure Support Department supporting cloud platform operations...SuggestedWeekly payTemporary workFlexible hours$65 per hour
...Experis Site Reliability Engineer III Pay Range $65 What's the Job? Design, build, and maintain automated, scalable cloud infrastructure solutions. Develop and implement Infrastructure-as-Code (IaC) using tools like Terraform, Ansible, and YAML. Support...SuggestedWeekly payTemporary workFlexible hours$60 - $65 per hour
...The Asset Management Group (AMG) is seeking a Site Rel Engineer Sr to support our Run the Bank (RTB) application portfolio within the Software Reliability & Controls (SRC) organization. This role is focused on production stability, operational excellence, and risk management...SuggestedContract workTemporary work$117k - $209.33k
...Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SuggestedFull timeFor contractorsRemote work$61k - $101k
...We require formal training or certification in infrastructure engineering concepts plus 3+ years of applied experience. We need strong... ...and programs may include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...Full time$61k - $101k
...Salary: $61,000 - 101,000 per year Requirements: Formal training or certification in site reliability engineering concepts, plus 3+ years of hands-on experience Experience supporting SRE practices for data management or migration platforms and products Familiarity...Full time- ...Job Description Job Description BCforward is currently seeking a highly motivated Site Reliability Engineer Job Title: Site Reliability Engineer Location: Columbus, OH/Plano, TX or Tampa, FL Duration: Contract to hire 4 months Job Description We are...Contract work
- ...tasks using scripting and tools Python Bash etc Collaborate with development infrastructure and support teams to improve system reliability Drive adoption of SRE practices like SLIs SLOs and error budgets Ensure performance optimisation capacity planning and...Permanent employmentTemporary workWork experience placement
$96.8k - $145.2k
...Company: NTT DATA Services We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX), United States (US). Job Responsibilities Own and manage observability using New Relic (APM, infrastructure monitoring, dashboards...Temporary workFlexible hours- ...Job Title Site Reliability Engineer About Skyhigh Security Skyhigh Security is a dynamic, fast‑paced, cloud company that is a leader in the security industry. Our mission is to protect the world’s data, and because of this, we live and breathe security. We value...Full timeFlexible hoursShift work
$123.4k - $222.53k
...employees! Ready grow your career as part of the Uncarrier journey at T-Mobile? Our team is searching for our next Sr. Site Reliability Engineer to strengthen the reliability and resilience of the systems powering T-Mobile's payment platforms, enabling faster, safer...Work experience placement- ...automate them. # Experience in Implementing AI/ML-based monitoring and self-healing solutions. # Experience in Implementing Chaos Engineering/testing. Seniority level Mid-Senior level Employment type Full-time Job function Consulting, Analyst, and...Full time
- ...Site Reliability Engineer Hybrid Onsite Worker is required to work onsite 2-3 days per week in Phoenix, AZ OR Plano, TX Main Responsibilities ~ Experience in leading observability initiatives as lead engineer. ~ Development and implementation of build release...Work experience placement2 days per week3 days per week
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure & Production Management sector of Consumer & Community...
$61k - $101k
...Salary: $61,000 - 101,000 per year Requirements: We expect formal training or certification in site reliability engineering, along with 5+ years of applied experience We need deep expertise in reliability, scalability, performance, security, enterprise architecture...Full time- ...Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Corporate and Investment...
- ...CloudOps— the team that keeps Splunk Cloud running for some of the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines at a scale very few teams ever get to operate at. When the...
$80k - $104k
...certification, or equivalent. Responsibilities: Lead reliability engineering initiatives across our Azure environment and Command Center... ...Architect Exposed Firewall More: We are hiring a Senior Site Reliability Engineer (Azure) for a six-month contract-to-...Full timeContract workShift workNight shift$107.48k - $143.31k
...employees. What You'll Be Doing Lead and apply regional reliability engineering strategies to improve equipment performance, uptime, and... ...management skills with the ability to support multiple sites remotely. Willingness to travel to the plant and corporate...Temporary workRemote workFlexible hours$105.9k - $145k
...are seeking an experienced Senior Software Engineer to design, develop, and modernize our... ...and collaborate across teams to deliver reliable and maintainable software solutions.This... ...operate. Visit our Corporate Sustainability site to learn more about our culture and commitment...Permanent employmentH1bWork at officeLocal areaRemote work$92.35k - $130k
.... We are seeking an experienced Software Engineer III to design, develop, and modernize our... ...and collaborate across teams to deliver reliable and maintainable software solutions.This... ...operate. Visit our Corporate Sustainability site to learn more about our culture and...Permanent employmentH1bWork at officeLocal areaRemote work- ...We are looking for a skilled Mobile Engineer to perform technical and mechanical functions related to the property's plant and equipment... ...inspections as assigned and completing them on time Conduct Site inspections to include roof, mechanical & electrical rooms as assigned...Hourly payH1bWork at officeLocal areaImmediate startVisa sponsorshipWork visaFlexible hours
- ...: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven environment...Long term contractContract workInternshipRemote work
$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work- ...where inclusion, sustainability, and community impact are more than values, they’re how we operate. Visit our Corporate Sustainability site to learn more about our culture and commitment to our people, customers, community, environment, and shareholders. Equal Employment...Full timeH1bWork at officeLocal areaRemote work
- ...GEICO is seeking a Senior SRE Software Engineer to design, build, and operate high-performance distributed platforms with zero-downtime reliability. You will own incident management tooling, lead on-call operations, and help scale automation across complex systems....
$95k - $115k
...Services is seeking a talented and experienced Mid-Level Software Engineer to join our growing Application Engineering department. This... ..., ensuring rapid features delivery and high system reliability across active product and T&M client queues. Roles & Responsibilities...Full timeWork at officeRemote workFlexible hours- ...GEICO is seeking an experienced Staff Engineer in the Technology Operations Center to lead incident management tooling and platform reliability. You will design, develop, and operate automation, dashboards, and data pipelines across incident response, on-call, and runbooks...
$83.8k - $107.6k
...Solutions is seeking an adaptable, high-impact Senior Software Engineer (AI Operations) to spearhead the operational stability,... ...Requirements Experience: 4+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), telecommunications operations, or mission-...WorldwideShift workNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Allen, TX
- construction site safety Allen, TX
- junior website developer Allen, TX
- official site Allen, TX
- junior site reliability engineer
- site reliability engineering manager
- site reliability engineer
- lead site reliability engineer
- site reliability engineer remote
- site reliability engineer sre



