Site Reliability Engineer
GrabJobs
Essential Functions: Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness. Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives. Evaluate system architecture and changes to ensure they balance functional requirements, service quality, reliability, security, and compliance needs. Drive continuous improvement in platform stability, maintenance, and availability. Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders. Experience and Skills Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or related infrastructure roles supporting production services. Strong experience with Linux systems administration and troubleshooting in enterprise environments. Strong experience operating and maintaining on-prem Kubernetes platforms and all related components including CRI, CNI, and CSI plugins. Experience deploying and maintaining applications on Kubernetes using Helm, Kustomize, and similar tooling. Experience supporting DevOps tooling such as GitLab, Artifactory, Jira, Confluence. Experience with GitOps tools such as FluxCD or ArgoCD. Proficiency scripting with at least one of Python, Go, or Bash. Strong experience designing, maintaining, and maturing observability tooling including monitoring, dashboards, logging and tracing, and supporting SLOs. Strong understanding of reliability engineering concepts: Service health indicators High availability design, failure reduction, and testing Operational readiness practices, including developing documentation, runbooks, and architectural descriptions Incident response, root cause analysis, remediation/recovery Ability to obtain a security clearance, which includes U.S. citizenship. Preferred: Experience with multiple Linux distributions including Ubuntu. Experience with at least one of the following: Tanzu Kubernetes, Nutanix Kubernetes Platform, Canonical Kubernetes. Experience with cloud platforms such as AWS and Azure. Experience with infrastructure automation and configuration management. Experience managing AI tooling on Kubernetes including MCP Servers, LLM platforms (vLLM, Ollama), Kubeflow. Experience with security and compliance considerations in regulated environments. DoD experience. Active or inactive Secret Security Clearance. Education: Bachelor’s degree in CS, Software Engineering or other IT-related field or equivalent experience REMOTE WORK NOTICE: This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be given to candidates located onsite in the Albuquerque area.
$125k - $150k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor...SuggestedPermanent employmentTemporary work$125k - $145k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of...SuggestedPermanent employmentTemporary workFlexible hoursWeekend work- ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus...SuggestedFull timeLocal areaWorldwideFlexible hours
$165k - $265k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...SuggestedPermanent employmentTemporary workWorldwideWeekend work$165k - $230k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts....SuggestedPermanent employmentTemporary workImmediate startWeekend work$165k - $265k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink...Permanent employmentTemporary workWorldwideWeekend work$145k - $175k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (TOP SECRET CLEARANCE)As a member of the Classified IT Systems Engineering team, the Site Reliability Engineer is involved...Permanent employmentTemporary workWeekend work$125k - $145k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCEAs a Site Reliability Engineer, you will design, develop, and test key aspects of an in-house...Permanent employmentTemporary workWeekend work$32 - $35 per hour
...and assignment.) Key Responsibilities: In this role, you will help ensure the reliability, performance, and stability of key restaurant-facing platforms by working closely with engineering and infrastructure teams. You will use observability tools such as DataDog, Grafana...Contract workLocal areaImmediate start$119k - $170k
...the greater good, come make your next move with Zscaler. Our Engineering team built the world’s largest cloud security platform from... ...cloud-first strategy. We’re looking for an experienced Staff Site Reliability Engineer (Federal) to join our Government Cloud team....Full timeWork at officeLocal areaWorldwideNight shift- ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates...Remote work
$180.5k - $236.91k
...Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team. Oscar is the... ...on your team's business and technical domains such as DevOps, site reliability, and cloud best practices Lead the planning, execution and...Full timeWork at officeRemote work$140k - $180k
...fundamentally different class of spacecraft. Engineered to survive the harshest radiation... ...create highly available, deployable, and reliable products Reduce operational toil through... ...experience in Software Engineering, Site Reliability Engineering or DevOps ~ Deep...Permanent employmentShift work- ...What you will do: Partner with a team of high-performing engineers and developers who are focused on delivering best in class software... ...our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability Building Zero trust...Full timeContract workLocal areaFlexible hoursShift work
- ...SRE Support Engineer While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (iNDIA)(8AM–5PM IST) Domain: Compute(Linux Fundamentals, Linux Networking, Kubernetes, Docker)...Remote work
$30.53 - $56.48 per hour
Job Title:Associate Site Reliability EngineerRequisition ID:R027696Job Description:Job Title: Associate Site Reliability EngineerReporting... ...TechnologyLocation: Santa Monica, CaOverviewThe Associate Site Reliability Engineer helps keep Marketing Technology services reliable, observable...Hourly payFull timeTemporary workPart timeInternshipLocal areaWorldwideRelocation package$164k - $270k
...for the 21st century and beyond.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to... ...remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop SLOs...Permanent employmentFull timeLocal areaFlexible hours$164k - $270k
Hadrian - Manufacturing the FutureHadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x cheaper. By combining advanced software, robotics, and full-stack manufacturing...Permanent employmentFull timeLocal areaRemote workFlexible hours- ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating... ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms...Permanent employmentFull timeTemporary workImmediate start
$197k - $291k
...troubleshooting distributed systems. Preferred qualifications Master's degree in Computer Science or Engineering. 1 year of people management experience. About The Job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale,...Full time$100k - $117.5k
...with the ultimate goal of enabling human life on Mars.BUILD RELIABILITY ENGINEER (FALCON/STARSHIP VALVES AND COMPONENTS) The Build Reliability... ...in quarantining and rework activities across multiple sites BASIC QUALIFICATIONS:Bachelor's degree in mechanical engineering...Permanent employmentTemporary workInternshipWork at officeImmediate startWeekend work$120k - $180k
...organizations that test and validate complex systems—think drones, rocket engines, satellites, and nuclear reactors. Supported by leading... ...to roll : Frequently traveling to spend time with end-users on-site (e.g. rocket test stands, spacecraft clean rooms, automated...Full timeTemporary workWork experience placement$120k - $145k
...the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER, SATELLITE SYSTEMS (STARSHIELD) Starshield leverages SpaceX’s... ...hosted payloads. The Starshield software team is building highly reliable in-space mesh networks, designing secure systems to guarantee...Permanent employmentFull timeTemporary workWeekend work- ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas... ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle...
- ...defense programs. Our platform gives hardware engineering teams a single place to ingest data,... ..., frequent travel to end-user sites, and requires U.S. TS Clearance eligibility... ...environments. Reduce complexity and improve reliability as we grow.Drive priorities: Identify the...Permanent employmentWork at office
$141.9k - $190.3k
...and transcends generations. We’re looking for passionate engineers who love learning new technologies at a rapid pace. You should... ..., and clear observability ~ Maintain and improve the reliability of services and infrastructure ~ Troubleshoot and resolve...Work experience placement- A leading technology company is seeking an Engineering Manager to lead a team focused on Site Reliability Engineering. The role demands a strong background in software development, data structures, and team management. Responsible for the uptime and performance of critical...
$120k - $150k
...love for you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleAs a Software Engineer, Business Systems you will have the opportunity to architect and manage the Apex “Operating System” platform from the ground up. Reporting...Full timeWork at office$155k - $205k
...to make this possible, with the ultimate goal of enabling human life on Mars.SR. DATABASE RELIABILITY ENGINEERSpaceX is looking for an experienced Database Reliability Engineer with expert-level technical knowledge and broad hands-on experience in OLTP as well as OLAP...Permanent employmentTemporary workFlexible hoursWeekend work$150k - $180k
WHAT YOU’LL DOThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess a deep understanding of systems engineering and automation, including configuration management...Work experience placementLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Los Angeles, CA
- website content developer Los Angeles, CA
- after school site coordinator Los Angeles, CA
- site leader Los Angeles, CA
- on-site clinical research associate (traveling/remote) Los Angeles, CA
- on site coordinator Los Angeles, CA
- official site Los Angeles, CA
- site recruiter Los Angeles, CA
- historic site Los Angeles, CA
- IT site lead Los Angeles, CA



