Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Inference Infrastructure

Cohere

Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.We are a global technology company co-headquartered in Toronto and San Francisco, with key offices in London, New York City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.As a Site Reliability Engineer you will:Build self-service systems that automate managing, deploying and operating services.This includes our custom Kubernetes operators that support language model deployments.Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.Develop our team through knowledge sharing and an active review process.You may be a good fit if you have:5+ years of engineering experience running production infrastructure at a large scaleExperience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clustersExperience with Kubernetes dev and production coding and supportExperience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid servingExperience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environmentsExperience in compute/storage/network resource and cost managementExcellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamworkThe grit and adaptability to solve complex technical challenges that evolve day to dayFamiliarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems.Experience in Golang, C++ or other languages designed for high-performance scalable servers).Full-Time Employees at Cohere enjoy these Perks:A weekly lunch stipend of $75/75 or equivalent in your local currency for lunch.Full health and dental benefits, including a separate budget for mental health.RRSP matching, 401K, Pension Scheme.100% Parental Leave top-up for up to 6 months, for either parent.Annual enrichment benefits:Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.Education & learning stipend for conferences, courses, and coaching.6 weeks of paid vacation (30 working days!)Budget for traveling to other offices if you are remote, plus an annual company offsite.How and Where We Work:Cohere is remote-friendly. We have offices in Toronto, San Francisco, New York City, London, Paris, Montreal, and more coming soon.For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.For those not near an office: a co-working benefit so you can work alongside others in your city.Everyone receives a $500 home office stipend to set up your workspace properly.If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.LocationToronto; London; Montreal; New York; San FranciscoEmployment TypeFull timeLocation TypeRemoteDepartmentInferenceModel Serving

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Inference Infrastructure in New York, NY vacancy
  • $139k - $257.55k

     ...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through...  ...AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and...  ...platform migrationBuild and operate ML inference infrastructure — model serving, GPU... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    8 hours ago
  • $300k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude...  ...providers, and make smart infrastructure decisions that keep us cost-... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    12 hours ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering...  ...the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI...  ...platform, balancing latency, throughput, reliability, and cost. Partner across... 
    Suggested
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    4 days ago
  • $152k - $241.5k

     ...built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving...  ...multi‑GPU inference performance and reliability: parallelism strategies,...  ...benchmarking and performance regression infrastructure for latency/throughput.Systems performance... 
    Suggested
    Full time
    Remote work

    Nvidia

    New York, NY
    1 day ago
  • $320k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...Our mandate is to make inference deployment boring and unattended...  ...design and build the deployment infrastructure that moves inference code... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    New York, NY
    12 hours ago
  • $300k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for building...  ...high-performance inference infrastructure they need to develop next-... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    12 hours ago
  •  ...About Us: Modal provides the infrastructure foundation for AI teams....  ...jobs, and serve low-latency inference. We have thousands of customers...  ...medalists, and experienced engineering and product leaders with...  ..., we seek to improve our reliability dramatically while scaling... 

    Modal Labs

    New York, NY
    4 days ago
  •  ...We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability...  ...highly available and fault-tolerant infrastructures to support our web services and ML...  ...• Make sure our platform, inference and model training environments are... 
    Relocation package

    Mistral Ai

    New York, NY
    2 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied...  ..., and frameworks—operating as an infrastructure-focused engineer, not a tool operator...  ...tooling/dashboards, including GPU/inference and token cost attribution).Prior software... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    New York, NY
    3 days ago
  • $140k - $170k

     ...optimize business interactions. Role Description: As a Site Reliability Engineer, you will work with Agile engineering teams to provide...  ...customer experiences, have deep experience with infrastructure, operational automation, data driven metrics collection,... 
    Full time
    Local area

    Symphony Communication Services

    New York, NY
    4 hours ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction... 
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    3 days ago
  • $140k - $225k

     ...on private equity, real estate, public debt and equity, infrastructure, life sciences, growth equity, opportunistic, non-investment...  ...on LinkedIn, X, and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability of systems... 
    Full time
    Local area
    Flexible hours

    Blackstone Group

    New York, NY
    4 days ago
  • $131k - $164k

    Position Overview We are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation frameworks, to join our global Infrastructure & Operations team. This role is a hands-on senior engineering position... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent Corporation

    New York, NY
    2 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For...  ...in technology infrastructure and world-class talent —...  ...available through this site. Capital One Financial... 
    Full time
    Part time
    Local area

    Capital One Financial Corp

    New York, NY
    3 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview...  ...creating responsible and reliable AI systems, changing...  ...investments in technology infrastructure and world-class talent — along...  ...information available through this site. Capital One Financial... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $141k - $216.6k

     ...are creating the only platform that combines modern 911 infrastructure with an AI intelligence layer—helping public safety agencies...  ...a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    2 days ago
  • $138.1k - $198.2k

     ...technology that simply works.  The SRE Engineering Enablement Team supports our CI...  ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter...  ...developer platform, ensuring that our infrastructure is not just functional, but a catalyst... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    New York, NY
    3 days ago
  • $165k - $241.4k

     ...aspects of the Federal region’s infrastructure and operations, such as...  ....We’re looking for talented engineers with a software or operations...  ...development teams to ensure the reliability, performance and security of...  ...see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    3 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team...  ...for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    8 hours ago
  • $120k - $200k

     ...DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal...  ...implementation of the company’s global infrastructure architecture. Responsible for cross-region...  ...testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    1 day ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    New York, NY
    4 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment...  ...and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications... 
    Shift work

    JP Morgan Chase

    New York, NY
    8 hours ago
  • $158.5k - $172k

     ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you...  ...design, automate, and secure our core infrastructure, deployment platforms, and CI/CD ecosystems...  ...-impact position driving continuous reliability, deep system optimization, and automation... 
    Full time
    Work at office
    3 days per week

    GrubHub

    New York, NY
    3 days ago
  • $110k - $120k

     ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a...  ...resolve complex production incidents across application and infrastructure layers.Act as the L3 escalation point for operational issues... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    8 hours ago
  • $123k - $165k

    Job Posting Title:Site Reliability Engineer IIReq ID:10143234Job Description:Department/Group OverviewOur engineering fleet is a horizontal...  ...strategies for distributed systems.Maintain and improve Infrastructure-as-Code (IaC) definitions and cloud environment configurations... 
    Full time

    Hulu

    New York, NY
    4 days ago
  • $194k - $267k

     ...potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era....  ...:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    3 days ago
  • As a Site Reliability Engineering at JPMorgan Chase within the Enterprise technology, liquidity risk team, you are the non-functional requirement owner and champion for the applications in your remit. You are a key influencer in your team’s strategic planning, driving... 

    JP Morgan Chase

    New York, NY
    3 days ago
  • $194k - $267k

     ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and...  ...that maximize platform reliability and engineering velocity.The ideal candidate is someone... 
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    2 days ago
  • Guide and shape the future of technology at a globally recognized firm, driven by pride in ownership.As a Senior Manager of Site Reliability Engineering at JPMorgan Chase within the Corporate Investment Bank, Markets team, you are the non-functional requirement owner and... 
    Bank staff
    Shift work

    JP Morgan Chase

    New York, NY
    1 day ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible...  ...and systems, architect low latency infrastructure solutions, proactively guard against...  ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Inference Infrastructure. Be the first to apply!