Site Reliability Engineer
Baseten
Site Reliability Engineer
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently.
You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.
You'll work on projects like these as part of the SRE team:
- Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services.
- Building AI-assisted tooling for incident triage and response.
Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code.
Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution.
Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations.
Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.
Define and instrument SLOs and SLIs across customer workloads and internal services.
Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.
- Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus).
- Experience in building and maintaining scalable infrastructure.
- Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines. Observability-as-code experience is a plus.
- Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).
- Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.
- Comfort working at the intersection of engineering and operations — you write code, but you also think deeply about process, escalation paths, and operational leverage.
- Familiarity with incident management platforms (incident.io or similar) is a plus.
- No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.
Competitive compensation, including meaningful equity
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
Paid parental leave
Fertility and family-building stipend through Carrot
Company-facilitated 401(k)
Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...SuggestedWork experience placement$45 - $85 per hour
DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software...SuggestedContract workTemporary work$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$141k - $216.6k
...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical...SuggestedWork experience placementWork at office$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Full timeTemporary workLocal areaRemote workWorldwide$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as a Site Reliability...Full time$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating...Full time- ...to meet you. Our Enterprise Information Technology (EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this role, you will move beyond traditional infrastructure maintenance...Permanent employmentFull timeH1bLocal areaRemote workShift work
$191k - $226k
...incentives to steer members to the care that helps them get healthier, faster. About the role We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML...Remote workWork visaFlexible hours- ...Morgan with security as a key differentiator.We pride ourselves on our "hands off culture" by having very few meetings and giving engineers and creators a broad scope of responsibility and autonomy. You'll be working in-person 5 days/week in our lower-manhattan NYC office...Full timeWork at office
$115k - $160k
...Senior Site Reliability Engineer - AVP - Credit Trade FloorEmbark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear – to redefine the future of banking and help craft innovative solutions. You...Work at office$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement...$180.5k - $236.91k
...Senior Software Engineer, Cloud Infrastructure / SRENew York, New York, United StatesHi, we're Oscar. We're hiring a Senior Software... ...your team's business and technical domains such as DevOps, site reliability, and cloud best practicesLead the planning, execution and release...Full timeWork at officeFlexible hours- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
- ...SRE Engineer Location: New York, NY, USA Experience: 8-12 Years Client: Amex Job Description: This is an SRE role supporting the B2B and Core Services. This is not a DevOps role, strictly need an SRE Engineer, who has great analytical skills and is a good...
- ...real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment. We’ve... ...rate limiting, and permissions scoping, etc. (gist of harness engineering) # You have experience building or orchestrating AI/agent workflows...Temporary workImmediate start
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area$160k - $230k
...Cloud AWS Support Reliability Engineer (SRE)When you work at the New York Fed, you have the opportunity to make an impact in our communities and across the nation. Our mission-driven, curious, and dedicated colleagues apply their diverse perspectives and unique talents...Full time- ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You...Full timeWork at office
- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...
$130k - $200k
...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership...Shift work- ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure...Full timeWork at officeRemote workFlexible hours2 days per week
$115k - $125k
...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help...Work experience placementWork at officeWork from homeMonday to FridayFlexible hoursShift workWeekend workAfternoon shiftEarly shift$115k - $160k
...Senior Site Reliability Engineer - AVP - Credit Trade FloorEmbark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear – to redefine the future of banking and help craft innovative solutions. You...Hourly payWork at office$105k - $300k
...Site Reliability Engineer At Citadel, a leading investor in the world's financial markets, we aim to win together as one team to earn the long-term trust of our capital partners and each other. Our collaborative approach allows technologists to grow alongside other...- ...production paths and high-volume event processing in a fast-moving startup environment. We are looking for a Platform/Infrastructure engineer with 4+ years in SRE or DevOps, strong GCP experience, and hands-on work with serverless systems, IAM, and observability. #J-188...
$150k - $300k
...What We Do At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people... ...Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,...Full timeTemporary workPart time- ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect...
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site recruiter New York, NY
- site services specialist New York, NY
- junior website developer New York, NY
- remote website tester New York, NY
- official site New York, NY
- on site coordinator New York, NY
- site leader New York, NY


