Site Reliability Engineer
GrabJobs
Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team. AI Studio is a part of Nebius Cloud , one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that promise, we need an engineer who can make the platform behave flawlessly under extreme load and recover gracefully when the unexpected happens. In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets. Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about. If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story. What we offer Competitive salary and comprehensive benefits package. Opportunities for professional growth within Nebius. Hybrid working arrangements. A dynamic and collaborative work environment that values initiative and innovation. We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!
- ...innovative ways to help people? Do you like having the autonomy to build new solutions from the ground up? If so, being a Software Engineer II at Frost could be the job for you.At Frost, it’s about more than a job. It’s about having a flourishing career where you can...SuggestedFull time
$89k - $110k
...the assigned supervisor for performing engineering work activities and projects requiring the... ..., Distribution Dispatch, Distribution Reliability, Distribution Planning, management and a... ...programWhere you’ll work:This role sits on site in New Albany, Ohio, Roanoke, Virginia,...SuggestedFull timeRemote workFlexible hours- ...cohesion. Supports department and team by engaging in patient transport as appropriate. Obtains and maintains proficiency in access site management. Acts as radiation safety representative for patients and team while X-ray is used. Performs fluoro imaging and...SuggestedFull time
$116.36k - $155.15k
...demonstrated knowledge and experience in system architecture and engineering disciplines. Specific technical knowledge of enterprise level... ...Amazon Web Services. Supports due diligence activities including site surveys, design, design review, bill of materials creation,...SuggestedFull timeTemporary workRemote work1 day per week$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting – Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...SuggestedWork at officeRemote work- ...technology and design. Overview The company is seeking a Senior Platform Engineer to join their Integrations Platform team. This role is responsible for the architecture, scalability, and reliability of the systems powering integrations, agents, and workflows across the...Permanent employmentRemote work
$150k - $220k
...d love to meet you. About the Role As a Senior Platform Engineer on the Data Platform team, you will be responsible for designing... ...the software development lifecycle. Drive security, reliability, scalability, observability, and cost-efficiency improvements...Full timeWork at officeLocal areaRemote work- ...develop models of possible future configurations Create daily test metrics and reporting Occasionally perform other IT systems engineering activities such as requirements, design, installation, operation, sustainment, and support Apply technical principles,...Full timeWork experience placementRemote work
$75.6k - $126k
...About Us Role Summary We are seeking an AI Solutions Engineer to guide the strategy, architecture, and hands-on development of AI-powered agents, workflows, and automation solutions that materially improve Sales performance. This role combines elements of AI product management...Work at officeLocal areaRemote work- ...Synera, we’re looking for a Forward Deployed Engineer (FDE) who enjoys solving meaningful... ...agents and multi-agent systems to make them reliable Communicate closely with Synera product... ...colleagues during regular team events and off-sites (two big off-sites in Germany). ABOUT US...Permanent employmentTemporary workWork experience placementWork at officeHome officeFlexible hours
$145k - $260k
...quickly. We are looking for outstanding product-focused software engineers who want to work on difficult technical problems with the goal... ...of all developers. Improve Warp's performance and reliability by building new features or polishing existing parts of the product...Work at officeImmediate startRemote workFlexible hours$140k - $160k
...of. About the Role WellTheory is looking for an Implementations Engineer to join our fully-remote team. This is a unique opportunity to... ...implementation — collaborating hands-on to get each one live quickly and reliably Make foundational technical decisions as part of a small,...Remote work$137.7k - $182.25k
...We’re hiring Senior Software Engineer I's to join the Integrations team within the Expansion product group! The mission of this team is... ...efficiency and effectiveness of our creator personas by delivering reliable integrations that embed Articulate 360 into our customers’...Local areaImmediate startRemote work$152.65k
...health equity. Who We're Looking For We are looking to hire a Software Engineer II, Integrations (L2) who will deliver scoped integration work that connects external systems to our platform in a reliable, maintainable, and secure way. You will work hands-on building and...Temporary workRemote workFlexible hours$153k - $179k
...cardiologists and are a growing team of world-class engineering, operations, medical affairs, marketing,... ...with some roles requiring you to be on-site in a location. Cleerly has created a new... ...-incident reviews, driving long-term reliability improvements across systems....Remote work- ...are scheduled. Generally these are scheduled between 1pm-6pm CET. About You Celestia Labs is looking for an elite Software Engineer to join the Celestia Node Team. You will be working on a highly technical team, operating across a cutting edge set of disciplines...Live inRemote workHome officeFlexible hours
- ...resilient EVP/TSP services. Requirements Requires Master's degree or foreign education equivalent in Computer Science or Computer Engineering + 3 years' experience in a software development role. Alternatively, Bachelor's degree + 5 years' experience. This is a...Remote workFlexible hours
- ...exposure to how a startup scales with visibility into product, engineering, design, and go-to-market. Thinking of founding a startup... ...great looks like for modern Rust SDKs—making them idiomatic, reliable, and delightful to use. You’ll own parts of our open source compiler...Work at officeRemote workWork from homeFlexible hours
- ...Redwire Defense Tech is seeking an experienced Autonomy/AI software engineer with a strong background in implementing optimization... ...operational aircraft engaged in global missions, requiring robust, reliable, and safety-focused software development. RESPONSIBILITIES The...Temporary workWork at officeRemote work
$154.45k - $185.33k
...NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly,... ...the U.S. The Opportunity Grafana Labs is seeking a Senior Engineer (AI & Automation) to own the AI agent infrastructure and automation...Local areaRemote workFlexible hours- ...Capital. About the Team and Role: We are hiring a senior software engineer to help design and build core components of our next-... ...significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components...Local areaWork from homeFlexible hours
$141k - $180k
...We're looking for talented and highly motivated software engineers to join our team. As a growth stage technology company, Clear Capital is seeking developers eager to be part of a team, working closely with every part of the company from operations to our executive team...Temporary workImmediate startRemote work$145k
...work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and... ...data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation...Full timeTemporary workWork at officeLocal areaRemote workFlexible hours$140k - $200k
...- Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and... ...→ testing → release → maintenance. Ensure quality, reliability, and consistency across releases. Identify, diagnose, and resolve...Work at office- ...SECURITY CLEARANCE: TS/SCI with Polygraphs required POSITION: Senior CNO Software Engineer LABOR CATEGORY: Senior Software Engineer REQUISITION: TKO-SWE3-06.232025 LOCATION: Ft. Meade, Maryland CNO is needed! CSG is seeking analytic development skills...Work experience placementRemote work
$140k - $200k
...people around the globe work on Speechify in a 100% distributed setting - Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups...Full timeWork at officeShift work- ...modernize secure, scalable solutions using cloud platforms and top engineering practices. Allata also empowers clients to unlock data... ...leverage LLMs as assistive components. This role focuses on creating reliable, explainable, and human‑validated solutions where AI enhances...
$141.52k - $213.9k
...made by real Twilions! . See yourself at Twilio Join the team as Twilio's next Senior or Staff Applied Software Research Engineer for Emerging Technologies. We are hiring for multiple positions at either a Senior or Staff level. Who we are At Twilio, we'...Local areaRemote workWorldwide- ...Senior Software Engineer, Services (this role is US based remote, all candidates must be authorized to work in the United States without... ...thinking and hands-on execution, with a focus on building reliable, extensible services that support multiple use cases over time....Remote work
$200k - $250k
...Senior Software Engineer, Transaction Management About Us Turnkey is developer-first infrastructure for private key management, making it simple to create wallets, sign transactions, and automate on-chain actions through one elegant API, without ever exposing sensitive...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site services specialist Corpus Christi, TX
- construction site safety Corpus Christi, TX
- site leader Corpus Christi, TX
- official site Corpus Christi, TX
- IT site lead Corpus Christi, TX
- site safety Corpus Christi, TX
- junior website developer Corpus Christi, TX
- on-site clinical research associate (traveling/remote) Corpus Christi, TX
- historic site Corpus Christi, TX
- site reliability engineering manager




