Site Reliability Engineer
GrabJobs
Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team. AI Studio is a part of Nebius Cloud , one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that promise, we need an engineer who can make the platform behave flawlessly under extreme load and recover gracefully when the unexpected happens. In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets. Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about. If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story. What we offer Competitive salary and comprehensive benefits package. Opportunities for professional growth within Nebius. Hybrid working arrangements. A dynamic and collaborative work environment that values initiative and innovation. We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!
$130k - $150k
...Site Reliability Engineer (SRE) Engineer Reliability into the Systems That Move the Nation’s Food Supply Who We Are US Cold owns and operates one of the most complex temperature-controlled logistics networks in North America. Every day, our systems coordinate...Suggested- ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering...SuggestedWork experience placementRemote work
$131k - $227.13k
...Description: The 1LMX MES COE is seeking an engineer who will own infrastructure‑as‑code, cloud platform, and reliability for the Apriso environment on AWS. This role blends full‑stack development, DevOps, and Site Reliability Engineering (SRE) practices to deliver...SuggestedFull timeTemporary workWork experience placementWork at officeRemote workRelocationFlexible hoursShift work3 days per week$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...SuggestedFull timeTemporary workWork experience placementRemote work- ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas... ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle...Suggested
- ...Job Details: Lead Site Reliability Engineer The Lead Site Reliability Engineer is a senior technical leader responsible for the reliability, availability, and operational excellence of a cloud-based infrastructure and distributed platform. This role owns uptime,...
$160.8k - $214.1k
...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield... .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the...Full timeTemporary workLocal areaRemote workFlexible hours- This position is on-site in PhiladelphiaAbout ProsciaProscia is revolutionizing pathology, the last major frontier in healthcare... ...what’s possible in medicine.About This PositionAs a Reliability Engineer reporting to the VP, Technical Operations, you will own the reliability...Work at officeShift work
$79.9k - $97.3k
...$79,900.00 to $97,300.00 Job Summary All Current - Software Engineer, AI Platform Hands-on individual contributor We're looking for... ...integrity). Continuously improve latency, cost, and reliability, including prompt caching and model-tier selection Job Requirements...Full timeLocal area$110.5k - $208.34k
...fostering innovation, integrity, and exemplifying the epitome of corporate responsibility. Your Mission is Ours.The Work Radar Systems Engineering within Rotary & Mission Systems (RMS) is seeking a candidate to support the design, integration, and test verification of high...Full timeContract workTemporary workWork at officeRemote workFlexible hours- OverviewLutron is seeking a seasoned engineering leader with deep expertise on software and embedded systems who is passionate about teaching the craft of engineering to junior and mid-career engineers as they create world-class, smart lighting and lighting control systems...ApprenticeshipWorldwide
- ...operations and capital project departments on all equipment reliability matters. You will drive the overall site goal of maximizing equipment performance through... ...of Bachelor degree in Mechanical or Electrical Engineering or similar field of study is required. Minimum...Local area
- ...and government agencies, a list that spans across the country.Job DescriptionThe Opportunity: CapTech is seeking a SaaS platform engineer to play a hands-on role in the engineering, administration, and evolution of our enterprise application ecosystem. This is not a traditional...Visa sponsorshipWork visa
$113.89k - $187.1k
...at Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual due... ...SummaryThe Comcast Cloud Team is seeking an OpenStack Platform Engineer to help design, test, optimize and deploy our multi-region private...Full timeWork at officeRemote workWorldwide$117.6k - $196.2k
...test, and deployment workflows to increase reliability, repeatability, and speed.Promote best... ..., artifact management, and release engineering.Security & ComplianceImplement and maintain... ...office, with an expectation of being on-site 50% of your working hours to support collaboration...Full timeWork at officeLocal areaRemote workFlexible hours- ...platform architecture with a focus on security, cost-efficiency, and operational excellence. The position requires collaboration with engineering, data, and product teams within an Agile environment to improve platform capabilities. Success is measured by the robustness,...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
- ...Time$130k - $150kJoin a growing technology-driven organization supporting mission-critical environments through modern platform engineering and cloud-native infrastructure. This full-time opportunity is ideal for an experienced Senior Platform Engineer who is passionate...Full timeFlexible hours
$1,000 per month
...technology and customer-centric solutions.OverviewAs a Senior Backend Engineer on the Trust Platform team, you'll play a pivotal role in... ...coding practices. Experience with designing resilient and reliable systems that meet high security and regulatory standards.Nice to...Temporary workWork at officeImmediate startRemote workFlexible hours$79.25k - $130.73k
OverviewAs a GIS subject matter expert, you’re a natural at identifying the right analysis tools for the problem at hand. Not only do you create innovative solutions, you talk about solutions in ways that get others excited about the power of GIS technology. Join an account...Local area$78.8k - $131.3k
Backend API’s and Integrations Software Engineer III Are you a Backend API developer looking to work for a mission driven global organization?About the role: We’re seeking a Software Engineer III with strong backend engineering skills to design, build, and operate distributed...Full timeContract workLocal areaWork from home$90k - $125k
...organization, innovation and creativity is within our DNA. Come help us make every talent moment Phenomenal!The Forward Deployment Engineering team plays an integral role in the end to end customer journey including implementation, customer upsells and renewalsIf you like...Full timeTemporary workWork at officeRelocation$79.25k - $130.73k
OverviewAs a GIS subject matter expert, you’re a natural at identifying the right analysis tools for the problem at hand. Not only do you create innovative solutions, you talk about solutions in ways that get others excited about the power of GIS technology. Join an account...$99k - $225k
...Sales Solutions EngineerThe Opportunity:We are seeking a Lead Sales Engineer with deep expertise in cyber defense products, such as defensive... ...our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is...Full timeContract workPart timeWork at officeLocal areaRemote work$125.9k - $148.1k
...ideas.We are seeking a highly skilled Full-Stack Senior Software Engineer to join our team. In this role, you will design, develop, and... ...what’s next.Job ResponsibilitiesDevelop and maintain a secure, reliable, and scalable, and efficient platform spanning back-end persistence...Full timeContract workLocal areaFlexible hours- Role SummaryNuuly is hiring a Senior Software Engineer to join our Technology team. Building performant scalable cloud service environments... ...incident response and post-mortem processes to improve system reliability. You should be adept at picking up new technologies and...
- Must Have Skills:6+ years of experience as a software engineer working with C++, Go, C, Python, Ansible, and BashWell versed in data structures and AlgorithmsNetworking experienceLinux software development experienceBackground in Cable/ telecommunications industryNice...Contract work
- ...cases, Comcast prefers to have employees on-site collaborating unless the team has been... ...ecosystem of diverse devices online quickly and reliably through cloud-to-cloud API and event-... ..., and operational excellenceMentor engineers and guide technical decision-making across...Full timeWork at officeRemote workWorldwide
- ...solutions for the consulting, architecture, engineering, project controls, procurement,... ...Travel Policy.You may visit construction sites and will be required to take site safety... ...to site safety rules.Must have access to reliable transportation.Must have the ability to...Night shiftDay shiftAfternoon shift
- ...stakeholders to deliver feasibility/complexity assessments and think critically about backlog prioritizationPractical knowledge of software engineering best practices, including Agile development methodologies, DevOps, DataOps, MLOps, containerization (e.g., Docker), and...ApprenticeshipEasy work
$293.9k - $406.8k
...the TeamYou will join Cisco’s Identity Engineering Group, a foundational organization responsible... ...domain, with a strong emphasis on reliability, interoperability, and long-term scalability... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workLocal areaRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site services specialist Philadelphia, PA
- construction site safety Philadelphia, PA
- site leader Philadelphia, PA
- official site Philadelphia, PA
- website content developer Philadelphia, PA
- on site coordinator Philadelphia, PA
- IT site lead Philadelphia, PA
- site safety Philadelphia, PA
- junior website developer Philadelphia, PA
- on-site clinical research associate (traveling/remote) Philadelphia, PA

