Senior Site Reliability Engineer
GrabJobs
Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team. AI Studio is a part of Nebius Cloud , one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that promise, we need an engineer who can make the platform behave flawlessly under extreme load and recover gracefully when the unexpected happens. In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets. Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about. If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story. What we offer Competitive salary and comprehensive benefits package. Opportunities for professional growth within Nebius. Hybrid working arrangements. A dynamic and collaborative work environment that values initiative and innovation. We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!
- ...The Home Depot is seeking a Senior Software Reliability Engineer to join the Platform Reliability Engineering team, ensuring the resilience, performance, and security of our enterprise Cloud Platform. You will mentor junior engineers, lead incident triage, root cause...Senior
- ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems... ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes...Senior
$139k - $160k
...architecture expertise, and facilitating integration of capabilities for experimentation, test, and deployment. The Role: As a Senior Software Engineer you will design, develop, and implement real-time software for RF sensor systems compliant with open architecture...SeniorFull timeWork experience placementLocal areaNight shift$132.23k - $176.31k
...shape the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem....SeniorFull timeTemporary workRemote work- ...navigating pricing, demand, and competitive complexity. As a Senior Full-Stack Software Engineer, you'll own end-to-end delivery across our analytics... ..., turning complex pricing and revenue problems into reliable, production-grade software. You'll partner closely with...Senior
- ...The Information Technology Senior Management Forum is seeking a Manager of Production Support... ...and a deep understanding of software engineering practices. Key responsibilities include managing teams, driving Site Reliability Engineering practices, and overseeing incident...
- Job description Snowflake SRE JD Your Role Accountabilities Primarily responsible for administrating Snowflake environments on AWS Identify, tune, and fix the performance issues on priority. Diagnose and troubleshoot Snowflake related errors and work with team to raise...
- ...We have an immediate need for a Senior Release Train Engineer for a contract assignment located in Carmel, Indiana . The Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train (ART) by steering it to success and navigating the complexity...SeniorContract workWork at officeImmediate start
- ...'re pioneering the future of warehouse automation with our innovative robotic and software technology. We're seeking a Senior Software Engineer experienced in C++ and C# who is passionate about developing high-quality, robust and secure software. The hiring team is...SeniorLocal area
- ...availability. • Automation Experience with Build/deployment, Software Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in automating manual processes using Python, Ruby, Unix Shell (bash,...Immediate start
$152.13k - $162.13k
...challenge the status-quo. Unum is changing, and we’re excited about what’s next. Join us. General Summary: Unum Group seeks Site Reliability Engineers in Atlanta, GA. Applicants who are interested in this position may apply at (Ref #66753) for consideration. Design,...Temporary workWork at officeRemote work- ...Technical Support Specialist In Site Reliability Engineering (Sre) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management. The...
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office- ...advances cures by helping the world's most important research sites do their best work. Our solutions are now used by over 30,00... ...What You'll Bring to the Team: We are seeking a Site Reliability Engineer (SRE) to join one of our Scrum teams and help ensure the...Work at office
- ...We are looking for an experienced and passionate Senior Software Engineer who cares about leading teams, delivering cutting-edge software, mentoring other people and leveraging your love for technology to drive digital excellence. This is a new role on a highly skilled...SeniorWork at officeLocal areaHome office
$140k
...their lives. About the Role: As a Stable Kernel Senior Software Engineer , you play an essential role in setting our portfolio of... ...you learn new technologies. Your knowledgeable practice, reliability, and consultative nature make you an engineer that stakeholders...SeniorFull timeContract workTemporary workVisa sponsorshipWork visaFlexible hours- ...DEPLOY has been retained to find a senior-level Software Engineer to join us on an hourly, flexible contract basis and help build the next generation of AI-first web and SaaS products. You'll own significant workstreams across a portfolio of active client engagements...SeniorHourly payContract workImmediate startRemote workFlexible hours
$119.85k - $162.15k
...the Problems That Matter Most FinQuery is looking for a Senior Software Engineer to join our Engineering team. The Senior Software... ...across services and interfaces, and guiding work through to reliable, production-quality outcomes. The position owns technical execution...SeniorContract workWork experience placementCasual workWork at officeImmediate startWork from homeWorldwideFlexible hours- ...Position Summary: Are you a seasoned engineer who loves building cloud-native applications... ...Services Limson Team is looking for a Senior Software Engineer to help spearhead the... ...testing of systems for accuracy, reliability and optimal performance Constructs various...SeniorWork at officeRemote workMonday to Friday
- ...A leading gaming company in Atlanta seeks a Senior Software Engineer. The ideal candidate will drive operational excellence in software development, utilizing their expertise in Java and cloud infrastructure to solve complex problems and create robust applications. Responsibilities...SeniorFull time
$101.5k - $169.1k
...Company Cox Automotive - USA Job Family Group Engineering / Product Development Job Profile Sr Release Train Engineer... ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures....SeniorWork at officeRemote workVisa sponsorshipFlexible hoursShift work$120k - $140k
...Full-time Description Sharetec is looking for a Senior Software Engineer to join our team! At Sharetec, we believe in a people-first... ...using Docker and contribute to CI/CD pipeline health and reliability Work with relational and cloud databases including PostgreSQL...SeniorFull timeRemote workNight shift- ...Pratt & Whitney India is seeking a Reliability Analyst to support Entry into Service (EIS) programs. You will lead dependability activities... ...service decisions. The role requires a bachelor’s degree in engineering/mathematics and 3–5 years of reliability experience, with...SeniorFull time
- Warner Bros. Discovery is seeking a Sr. Principal Data Scientist based in Atlanta, Georgia. In this high-impact role, you will design and deliver advanced AI systems that drive key business decisions. We are looking for a candidate with 15–18+ years of experience in Data...SeniorFull time
- ...are we looking for? We’re looking for experienced Software Engineers who are passionate about building rock-solid payment products... ...You’ll know you’re the right candidate when you enjoy crafting reliable SaaS APIs and are also passionate about helping other developers...SeniorFull timeWork experience placementFlexible hours
$165k - $247.5k
...Senior Principal Software Engineer - AI Governance Atlanta, Georgia Strength in Trust OneTrust’s mission is to enable innovation through... ...Governance platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern...SeniorFull timeWork experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week- ROI is seeking a Principal Workday Consultant, Human Capital Management, to lead strategic advisory and solution architecture for healthcare clients. You will own the design and delivery of Workday HCM engagements, ranging from net-new deployments to complex tenant optimizations...SeniorFull time
- ...Human Services (DHS), Office of Information Technology, is seeking a qualified candidate for a contractor staffing position for a Senior Software Developer on Georgias Child Welfare technical team in Atlanta, Georgia. Complete Description: Job Responsibilities...SeniorFull timeContract workFor contractorsWork at office
$158.08k - $176.8k
...Senior SAP Functional Consultant (P2P, Tax, Vertex, Ariba, S/4) Location: Irvine, CA; Atlanta, GA; Columbus, OH (options available)... ...years SAP P2P and tax configuration experience. ~ Vertex Tax Engine expertise required. ~ Strong knowledge of SAP data structures...SeniorFull time- A leading AI research accelerator is seeking a software engineer to evaluate and refine AI-generated code. Candidates must have over... ...designing verification mechanisms, and ensuring code efficiency and reliability. This is a contractor position requiring flexible engagement...SeniorContract workFor contractorsRemote work10 hours per weekFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- sr hr business partner Atlanta, GA
- senior lighting artist Atlanta, GA
- senior planner Atlanta, GA
- senior hvac project manager Atlanta, GA
- senior technical product manager Atlanta, GA
- senior cloud infrastructure engineer Atlanta, GA
- senior etl developer Atlanta, GA
- senior wealth advisor Atlanta, GA














