Site Reliability Engineer, Inference Infrastructure
Cohere
Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!Why this role?Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.As a Site Reliability Engineer you will:Build self-service systems that automate managing, deploying and operating services.This includes our custom Kubernetes operators that support language model deployments.Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.Develop our team through knowledge sharing and an active review process.You may be a good fit if you have:5+ years of engineering experience running production infrastructure at a large scaleExperience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clustersExperience with Kubernetes dev and production coding and supportExperience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid servingExperience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environmentsExperience in compute/storage/network resource and cost managementExcellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamworkThe grit and adaptability to solve complex technical challenges that evolve day to dayFamiliarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems.Experience in Golang, C++ or other languages designed for high-performance scalable servers).Full-Time Employees at Cohere enjoy these Perks:A weekly lunch stipend of $75/75 or equivalent in your local currency for lunch.Full health and dental benefits, including a separate budget for mental health.RRSP matching, 401K, Pension Scheme.100% Parental Leave top-up for up to 6 months, for either parent.Annual enrichment benefits:Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.Education & learning stipend for conferences, courses, and coaching.6 weeks of paid vacation (30 working days!)Budget for traveling to other offices if you are remote, plus an annual company offsite.How and Where We Work:Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.For those not near an office: a co-working benefit so you can work alongside others in your city.Everyone receives a $500 home office stipend to set up your workspace properly.If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.LocationToronto; London; Montreal; New York; San FranciscoEmployment TypeFull timeLocation TypeRemoteDepartmentInferenceModel Serving
$300k
...Anthropic’s mission is to create reliable, interpretable, and... ...group of committed researchers, engineers, policy experts, and business... ...the Role The Cloud Inference team scales and optimizes Claude... ...providers, and make smart infrastructure decisions that keep us cost-...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours- Observable Intuition is seeking a Founding Infrastructure Engineer to define and own the production inference platform behind a new AI layer. You’ll build core systems... ...while ensuring scalability, security, and reliability across cloud and air-gapped deployments. You’ll...Suggested
$300k
...Anthropic’s mission is to create reliable, interpretable, and... ...group of committed researchers, engineers, policy experts, and business... ...About the role Our Inference team is responsible for building... ...high-performance inference infrastructure they need to develop next-...SuggestedFull timeWork at officeWorldwideVisa sponsorshipFlexible hours$320k
...Anthropic’s mission is to create reliable, interpretable, and... ...group of committed researchers, engineers, policy experts, and business... ...Our mandate is to make inference deployment boring and unattended... ...design and build the deployment infrastructure that moves inference code...SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and... ...platform migrationBuild and operate ML inference infrastructure — model serving, GPU...SuggestedFull timeTemporary workLocal areaRemote workWorldwide$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering... ...the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI... ...platform, balancing latency, throughput, reliability, and cost. Partner across...Temporary workFor contractorsWork experience placement$160k - $240k
Bloomberg L.P. in New York is seeking a Senior Software Engineer for AI Inference to design and build scalable infrastructure for machine learning applications. The ideal candidate will have over 5 years of software engineering experience, expertise in distributed systems...$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and... ...platform migration Build and operate ML inference infrastructure - model serving, GPU...Temporary workLocal areaRemote workWorldwide- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay... ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...Flexible hours
- ...We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability... ...highly available and fault-tolerant infrastructures to support our web services and ML... ...• Make sure our platform, inference and model training environments are...Relocation package
$158.1k - $213.8k
...experience- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of Object Oriented Design experience...InternshipWorldwideFlexible hours$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied... ..., and frameworks—operating as an infrastructure-focused engineer, not a tool operator... ...tooling/dashboards, including GPU/inference and token cost attribution).Prior software...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...election libraries, and building infrastructure‑as‑code tooling that eliminated manual...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week$250k - $300k
Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction...Work experience placementWork at officeLocal areaImmediate start$140k - $225k
...on private equity, real estate, public debt and equity, infrastructure, life sciences, growth equity, opportunistic, non-investment... ...on LinkedIn, X, and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability of systems...Full timeLocal areaFlexible hours$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview... ...creating responsible and reliable AI systems, changing... ...investments in technology infrastructure and world-class talent — along... ...information available through this site. Capital One Financial...Full timePart timeLocal area$229.9k - $262.4k
...Sr. Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are... ...creating responsible and reliable AI systems, changing banking... ...in technology infrastructure and world-class talent —... ...information available through this site. Capital One Financial...Full timePart timeLocal area$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands‑on technically while also mentoring a small team of SREs. The InfraSec team collaborates...Local areaRemote workFlexible hours- Basis is seeking a Site Reliability Engineer to ensure the reliability, scalability, and performance of our AI-powered accounting platform. You’ll join a high-leverage infrastructure team at the intersection of product and platform, owning systems that keepBasis fast, secure...
$160k - $180k
Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience...- ...day, members of our team are hands-on scaling-out production infrastructure, building out CI/CD pipelines, and brainstorming with developers... ...and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS...For contractors
$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure...Work at officeLocal areaImmediate start$110k - $120k
...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a... ...resolve complex production incidents across application and infrastructure layers.Act as the L3 escalation point for operational issues...Ongoing contractFull timeCasual workRemote workFlexible hours$158.5k - $172k
...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you... ...design, automate, and secure our core infrastructure, deployment platforms, and CI/CD ecosystems... ...-impact position driving continuous reliability, deep system optimization, and automation...Full timeTemporary workWork at officeFlexible hours3 days per week$167.7k - $245.2k
...aspects of the Federal region’s infrastructure and operations, such as... ....We’re looking for talented engineers with a software or operations... ...development teams to ensure the reliability, performance and security of... ...see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$141k - $216.6k
...are creating the only platform that combines modern 911 infrastructure with an AI intelligence layer—helping public safety agencies... ...a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational...Work experience placementWork at office$120k - $200k
...DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal... ...implementation of the company’s global infrastructure architecture. Responsible for cross-region... ...testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)Overseas- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment... ...and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications...Shift work
- A leading software development company in New York is seeking an entry-level Python engineer to join their team in the Brooklyn office. The role involves working on the AI inference pipeline that powers sophisticated OCR and computer vision products. Candidates should have...Full timeWork at office
$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...and systems, architect low latency infrastructure solutions, proactively guard against... ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability...Full timeTemporary workPart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, Inference Infrastructure. Be the first to apply!
- site reliability engineer remote New York, NY
- site reliability engineer New York, NY
- site reliability engineer sre New York, NY
- data infrastructure engineer New York, NY
- infrastructure engineering manager New York, NY
- senior infrastructure engineer New York, NY
- infrastructure automation engineer New York, NY
- principal infrastructure engineer New York, NY
- remote infrastructure engineer New York, NY
- lead infrastructure engineer New York, NY


