Site Reliability Engineer
GrabJobs
Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team. AI Studio is a part of Nebius Cloud , one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that promise, we need an engineer who can make the platform behave flawlessly under extreme load and recover gracefully when the unexpected happens. In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets. Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about. If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story. What we offer Competitive salary and comprehensive benefits package. Opportunities for professional growth within Nebius. Hybrid working arrangements. A dynamic and collaborative work environment that values initiative and innovation. We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!
$160k - $210k
...departmental collaboration and a unified sense of purpose, making teamwork a cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and improve service management across Cognitiv. Our immediate...SuggestedWork at officeLocal areaImmediate startRemote work$210k - $220k
...secure and private by design, it’s popular with security, IT, engineering, finance, and other security-focused teams. At Tines, we're... ...we’re looking for others to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team responsible...SuggestedWork at officeRemote work- ...Site Reliability/DevOps Engineer As a Site Reliability Engineer, you will be a member of a cross-functional Engineering team building solutions to meet the needs of our customers, partners and internal operations staff. Your ability to lead your teammates towards robust...Suggested
- ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that...Suggested
$100 per hour
...our New York City HQ (EST) Where you'll create impact Improve reliability of our systems Build & maintain our main infrastructure (... ...learn, are highly curious about new frameworks and solutions to engineering problems Fast-moving: you deploy daily, iterate quickly, and...SuggestedImmediate startRemote workWork from homeRelocationHome officeVisa sponsorshipWeekend work$1,500 per month
...game worlds they inhabit. Our approach is centered around World Engine, our state-of-the-art onchain game server framework. World... ...architecture to keep our platform secure. Own delivery, scalability, and reliability of our backend infrastructure. Advise and collaborate with the...Full timeFlexible hours$148k - $193k
...Staff Site Reliability Engineer Denver, CO, USA DAT is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of innovation in transportation supply chain logistics for 45 years. We continue to transform...Temporary workWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...strategic Chainlink Reserve. Learn more at chain.link. The Engineering Team As adoption of the Chainlink Runtime Environment (CRE) accelerates, you will be a part of that growth to ensure reliability and security remain at the forefront of development. This role...Remote work
$87k - $105k
...bottlenecks, and improve system health—utilization, performance, and reliability—across our infrastructure Understand how systems fail and... ..., documentation, and code review—that automates reliability engineering work: Deployment tooling Fault-injection/chaos...Full timeCasual workLocal areaWorldwideFlexible hoursShift work$120k - $200k
...Coinbase Ventures, Uniswap Labs, Circle Ventures, Delphi Digital, and many more. ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering, dedicated to crafting and maintaining large-scale...Full time- ...Products Group, we are dedicated to excellence in the design and engineering of Lam's etch and deposition products. We drive innovation to... ...teams to design and develop software programs.May visit customer site to provide support and have ability to travel (total is less...Local areaRemote workFlexible hours2 days per week3 days per week1 day per week
$160.8k - $214.1k
...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield... .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the...Full timeTemporary workLocal areaRemote workFlexible hours- ...escalate risks, threshold concerns, and recurring issues to senior engineers or the Platform Manager.Participate in incident response as a... ..., and support knowledge articles.Schedule & Presence: This on-site role supports 24/7 operations through real-time collaboration,...Monday to FridayShift work
$136k - $184k
Platform Infrastructure Engineer IV or V, DOEHybrid (Office 3 days/wk - Onsite-Flex) within Portland, OR; Medford, OR; Renton, WA; Burlington... ...platform expertise in cloud, networking, automation, and reliability engineering—tackling complex technical challenges, influencing...Full timeWork at officeImmediate startWork from homeFlexible hours$115k - $135k
...meetingsWhat will you bring?Familiarity with or hands on experience using the Smarsh suite of products, OR 1-2 years of solution engineering experienceAn ability to demonstrate our products in a way to sells the value of our solutions aligned to a client’s business objectives...Full timeLocal areaRemote work- ...impact you’ll makeWe are seeking a Senior Engineer with dual role and responsibilities1.... ...technical deployments with the performance and reliability our customers expect, enabling fully... ...hybrid roles combine the benefits of on-site collaboration with colleagues and the...Local areaRemote workFlexible hours2 days per week3 days per week1 day per week
$139k - $249.26k
...Team, we have an exciting opportunity for an exceptional software engineer and subject matter expert to join our growing team, directly... ...Consulting M&E Team and our customersOccasional travel to customer sites and Autodesk offices may be requiredMinimum Qualifications5+...Full timeFor contractorsRemote work$48.27 - $76.22 per hour
Software Engineer II - RemoteThe Software Engineer is responsible for the development, implementation and maintenance of the information systems (IS) and technologies needed for the management and analysis of research and clinical trials data. They contribute to the strategic...Minimum wageLocal areaRemote workShift work- ...’ll be working with a team of other peer engineers as part of Nike’s Consumer Product & Innovation... ...and interpersonal conflict.Have reliable instincts regarding the consequences of complexity... ...have a rotating on-call schedule for off-site, local business-hours, contact. The team...Full timeLocal area
$98.18k - $115.5k
...institutions. We’re looking for people who want more than just a job - they want to make a difference! U.S. Bank is seeking a Software Engineer who will contribute toward the success of our technology initiatives in our digital transformation journey.This position will be...Full timeLocal area- ...Products Group, we are dedicated to excellence in the design and engineering of Lam's etch and deposition products. We drive innovation to... ...needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work...Local areaRemote workFlexible hours2 days per week3 days per week1 day per week
$126k - $158k
...for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the... ...top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally...Work at officeRemote workFlexible hours$143.7k - $194.4k
...You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers... ...of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years...InternshipWorldwideFlexible hours$119.77k - $140.9k
...adhering to architectural best practices; considers scalability, reliability and performance of systems/contexts affected when defining... ...meet standardsConducts code reviews to provide guidance on engineering best practices and compliance with development proceduresAccountable...Full timeWork experience placementLocal area3 days per week$293.9k - $406.8k
...the TeamYou will join Cisco’s Identity Engineering Group, a foundational organization responsible... ...domain, with a strong emphasis on reliability, interoperability, and long-term scalability... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workLocal areaRemote workFlexible hours- ...on consumer data that meets global expectations of security, reliability, and performance; collaborate with cross-functional teams to deliver... ...of post-baccalaureate experience in the job offered or a engineer-related occupation. Position requires: • Software Development...Remote work
$115k - $145k
...Portland / AtlantaDivisions - Corporate Engineering /Full-Time /RemoteAs a Software Engineer... ...meeting their compliance needs by enabling reliable, scalable, and high-performance search... ...Product Management, Engineering, and Site Reliability to solve complex challenges...Full timeWork at officeLocal areaRemote workFlexible hours$183.8k - $263.6k
...orchestration, and secure service integration. You will work closely with engineers across control plane, data plane, and platform teams to deliver... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaRemote workFlexible hours$97.02k - $163.03k
...development. We need the aid of a highly motivated problem solver. An engineer who enjoys the challenge of resolving complex problems with... ...analysis, design, peer reviews and documentationProvide reliable estimates of effort and good identification of risksApply sound...Minimum wageFull timeLocal areaRemote workFlexible hours$143.7k - $194.4k
...just getting started.We're looking for a Software Development Engineer to build the server-side systems that power Kiro Web — a local... ...years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years...InternshipLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site services specialist Portland, OR
- construction site safety Portland, OR
- site leader Portland, OR
- official site Portland, OR
- website content developer Portland, OR
- on site coordinator Portland, OR
- IT site lead Portland, OR
- site safety Portland, OR
- junior website developer Portland, OR
- on-site clinical research associate (traveling/remote) Portland, OR

