Site Reliability Engineer
GrabJobs
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding of Linux operating system internals and networking concepts (e.g., filesystems, TCP/IP, DNS, TLS) Possess a customer focused mindset, treating internal developers as your primary users Have strong operational ownership, including a track record of debugging complex production issues and driving them to resolution Prefer automation over manual processes ("allergic to ops work") We are a small team of software engineers with a strong bias toward building software solutions to eliminate toil Strong candidates may also have experience with Designing and implementing secure, multi-tenant runtime environments from first principles Proficiency with Kubernetes ecosystem tools such as Helm, Kustomize, Gatekeeper, Kyverno, and CRDs/Operators, CRI, CSI Expertise in cloud infrastructure platforms, including AWS, GCP, or Azure Proficiency in provisioning infrastructure using tools like Terraform, Crossplane, and AWS Controllers for Kubernetes (ACK) Advanced Linux systems internals and networking concepts specifically relevant to containers, such as namespaces and cgroups About MongoDB MongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the database for the AI era, enabling innovators to create, transform, and disrupt industries with software. MongoDB’s unified database platform, the most widely available, globally distributed database on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, MongoDB Atlas, is the only globally distributed, multi-cloud database and is available across AWS, Google Cloud, and Microsoft Azure. With offices worldwide and over 60,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on MongoDB for their most important applications, we’re powering the next era of software. Our compass at MongoDB is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us MongoDB. To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy , we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at MongoDB , and help us make an impact on the world! MongoDB is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter. MongoDB, Inc. provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type and makes all hiring decisions without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Req ID: 426182
- ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus...SuggestedFull timeLocal areaWorldwideFlexible hours
$86k - $148k
...make a difference. Position Summary We’re looking for a Senior Engineer to lead complex initiatives and elevate our managed services... ...incidents, implement monitoring solutions, and improve system reliability. Security-First Mindset: Experienced in aligning engineering solutions...SuggestedWork at officeImmediate startFlexible hours- ...Partner with software developers, platform engineers, and IT staff to improve system design,... ...requirements, service quality, reliability, security, and compliance needs. Drive continuous... ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform...SuggestedWork at officeRemote work
$32 - $35 per hour
...and assignment.) Key Responsibilities: In this role, you will help ensure the reliability, performance, and stability of key restaurant-facing platforms by working closely with engineering and infrastructure teams. You will use observability tools such as DataDog, Grafana...SuggestedContract workLocal areaImmediate start$119k - $170k
...the greater good, come make your next move with Zscaler. Our Engineering team built the world’s largest cloud security platform from... ...cloud-first strategy. We’re looking for an experienced Staff Site Reliability Engineer (Federal) to join our Government Cloud team....SuggestedFull timeWork at officeLocal areaWorldwideNight shift$180.5k - $236.91k
...Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team. Oscar is the... ...on your team's business and technical domains such as DevOps, site reliability, and cloud best practices Lead the planning, execution and...Full timeWork at officeRemote work- ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates...Remote work
- ...SRE Support Engineer While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (iNDIA)(8AM–5PM IST) Domain: Compute(Linux Fundamentals, Linux Networking, Kubernetes, Docker)...Remote work
$164k - $270k
...for the 21st century and beyond.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to... ...remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop SLOs...Permanent employmentFull timeLocal areaFlexible hours$164k - $270k
Hadrian - Manufacturing the FutureHadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x cheaper. By combining advanced software, robotics, and full-stack manufacturing...Permanent employmentFull timeLocal areaRemote workFlexible hours$140k - $180k
...fundamentally different class of spacecraft. Engineered to survive the harshest radiation... ...create highly available, deployable, and reliable products Reduce operational toil through... ...experience in Software Engineering, Site Reliability Engineering or DevOps ~ Deep...Permanent employmentShift work- ...What you will do: Partner with a team of high-performing engineers and developers who are focused on delivering best in class software... ...our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability Building Zero trust...Full timeContract workLocal areaFlexible hoursShift work
- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...InternshipWork at officeLocal areaRemote workWorldwide
$197k - $291k
...troubleshooting distributed systems. Preferred qualifications Master's degree in Computer Science or Engineering. 1 year of people management experience. About The Job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale,...Full time$120k - $180k
...organizations that test and validate complex systems—think drones, rocket engines, satellites, and nuclear reactors. Supported by leading... ...to roll : Frequently traveling to spend time with end-users on-site (e.g. rocket test stands, spacecraft clean rooms, automated...Full timeTemporary workWork experience placement- ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating... ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms...Permanent employmentFull timeTemporary workImmediate start
$141.9k - $190.3k
...and transcends generations. We’re looking for passionate engineers who love learning new technologies at a rapid pace. You should... ..., and clear observability ~ Maintain and improve the reliability of services and infrastructure ~ Troubleshoot and resolve...Work experience placement- A leading technology company is seeking an Engineering Manager to lead a team focused on Site Reliability Engineering. The role demands a strong background in software development, data structures, and team management. Responsible for the uptime and performance of critical...
- ...defense programs. Our platform gives hardware engineering teams a single place to ingest data,... ..., frequent travel to end-user sites, and requires U.S. TS Clearance eligibility... ...environments. Reduce complexity and improve reliability as we grow.Drive priorities: Identify the...Permanent employmentWork at office
$120k - $150k
...love for you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleAs a Software Engineer, Business Systems you will have the opportunity to architect and manage the Apex “Operating System” platform from the ground up. Reporting...Full timeWork at office$150k - $180k
WHAT YOU’LL DOThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess a deep understanding of systems engineering and automation, including configuration management...Work experience placementLocal area$100k - $200k
...backed by top tier investors. Our lean, world-class team of engineers and operators is applying a first-principles approach to... ...About This Role We are seeking a highly capable DevOps / Site Reliability Engineer to help build and operate the software systems underpinning...Full timeWeekend work- ...next-generation defense programs. Our platform gives hardware engineering teams a single place to ingest data, analyze performance,... ...CI/CD pipeline performance, release automation, and release reliability.Own developer infrastructure that enables engineering velocity...Permanent employmentWork at officeLocal area
$229.2k - $319.5k
...Principal Software Engineer - Developer Connections, Game Release Job Id: REQ-0010111 Riot engineers bring deep knowledge of specific... ...across Game Studios, R&D, and Central Tech to help teams reliably and efficiently ship and operate games. Our mission is to provide...Temporary workLocal areaImmediate startFlexible hours$155.9k - $233.9k
...excellence and creativity.SIE Studio IT is seeking a Lead Systems Engineer to lead the design, operation, and evolution of our production... ...by Systems Administrators to ensure quality, consistency, and reliability.Act as the technical lead during infrastructure incidents,...Shift workAfternoon shift- ...tier investors and has over $13M in government funding. About the Role Antares is seeking a Research & Development Software Engineer to build the software systems that enable fast, rigorous experimental engineering across our R&D organization. This role supports...Permanent employmentFull time
$200k - $250k
...way to offer a ticket to the millions of fans who browse our platform around the world. StubHub is seeking Senior Software Engineers to design and develop next-generation technologies and complex features that transform the way millions of users explore, interact...Full timeWork at officeRemote workWorldwideFlexible hours- ...Software System Engineer Here at The Exploration Company, we are developing, producing, and operating Nyx, a modular and reusable space orbital vehicle that can eventually be refuelled in orbit and that can carry cargo - and potentially humans in the longer run....Immediate startRelocationVisa sponsorshipRelocation package
$160k - $200k
...new era demands a fundamentally different class of spacecraft. Engineered to survive the harshest radiation environments and to fully... ...attitude control systems, and power systems to ensure safe and reliable operation of the vehicle. In your first 6 months you will developcore...Permanent employmentFull timeShift work$150k - $300k
...focused on supply chain connectivity. We're looking for a Software Engineer with 3-8 years of experience who brings strong technical... ...and engineering culture from day one. Location This is an on-site role based in Los Angeles, CA. Candidates should be local or willing...Full timeLocal areaRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site leader Glendale, CA
- on-site clinical research associate (traveling/remote) Glendale, CA
- official site Glendale, CA
- historic site Glendale, CA
- IT site lead Glendale, CA
- junior website developer Glendale, CA
- site safety Glendale, CA
- site services specialist Glendale, CA
- construction site safety Glendale, CA
- site reliability engineer




