Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Inference Infrastructure

Cohere

Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!Why this role?Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs.As a Site Reliability Engineer you will:Build self-service systems that automate managing, deploying and operating services.This includes our custom Kubernetes operators that support language model deployments.Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.Develop our team through knowledge sharing and an active review process.You may be a good fit if you have:5+ years of engineering experience running production infrastructure at a large scaleExperience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clustersExperience with Kubernetes dev and production coding and supportExperience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid servingExperience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environmentsExperience in compute/storage/network resource and cost managementExcellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamworkThe grit and adaptability to solve complex technical challenges that evolve day to dayFamiliarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems.Experience in Golang, C++ or other languages designed for high-performance scalable servers).Full-Time Employees at Cohere enjoy these Perks:A weekly lunch stipend of $75/75 or equivalent in your local currency for lunch.Full health and dental benefits, including a separate budget for mental health.RRSP matching, 401K, Pension Scheme.100% Parental Leave top-up for up to 6 months, for either parent.Annual enrichment benefits:Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.Education & learning stipend for conferences, courses, and coaching.6 weeks of paid vacation (30 working days!)Budget for traveling to other offices if you are remote, plus an annual company offsite.How and Where We Work:Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.For those not near an office: a co-working benefit so you can work alongside others in your city.Everyone receives a $500 home office stipend to set up your workspace properly.If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.LocationToronto; London; Montreal; New York; San FranciscoEmployment TypeFull timeLocation TypeRemoteDepartmentInferenceModel Serving

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Inference Infrastructure in New York, NY vacancy
  • $300k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude...  ...providers, and make smart infrastructure decisions that keep us cost-... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    1 day ago
  • Observable Intuition is seeking a Founding Infrastructure Engineer to define and own the production inference platform behind a new AI layer. You’ll build core systems...  ...while ensuring scalability, security, and reliability across cloud and air-gapped deployments. You’ll... 
    Suggested

    Observable Intuition

    New York, NY
    4 days ago
  • $300k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for building...  ...high-performance inference infrastructure they need to develop next-... 
    Suggested
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    15 hours ago
  • $320k

     ...Anthropic’s mission is to create reliable, interpretable, and...  ...group of committed researchers, engineers, policy experts, and business...  ...Our mandate is to make inference deployment boring and unattended...  ...design and build the deployment infrastructure that moves inference code... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    New York, NY
    15 hours ago
  • $139k - $257.55k

     ...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through...  ...AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and...  ...platform migrationBuild and operate ML inference infrastructure — model serving, GPU... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    1 day ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering...  ...the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI...  ...platform, balancing latency, throughput, reliability, and cost. Partner across... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    4 days ago
  • $160k - $240k

    Bloomberg L.P. in New York is seeking a Senior Software Engineer for AI Inference to design and build scalable infrastructure for machine learning applications. The ideal candidate will have over 5 years of software engineering experience, expertise in distributed systems... 

    Bloomberg

    New York, NY
    2 days ago
  • $139k - $257.55k

     ...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through...  ...AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and...  ...platform migration Build and operate ML inference infrastructure - model serving, GPU... 
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe

    New York, NY
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay...  ...Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... 
    Flexible hours

    Baseten

    New York, NY
    3 days ago
  •  ...We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability...  ...highly available and fault-tolerant infrastructures to support our web services and ML...  ...• Make sure our platform, inference and model training environments are... 
    Relocation package

    Mistral AI

    New York, NY
    1 day ago
  • $158.1k - $213.8k

     ...experience- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of Object Oriented Design experience... 
    Internship
    Worldwide
    Flexible hours

    Amazon

    New York, NY
    2 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied...  ..., and frameworks—operating as an infrastructure-focused engineer, not a tool operator...  ...tooling/dashboards, including GPU/inference and token cost attribution).Prior software... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    New York, NY
    4 days ago
  • $140k - $215k

     ...intersection of our Core Platform and Embedded Reliability charters: building the foundational...  ...while embedding directly with product engineering teams and their leadership to drive...  ...election libraries, and building infrastructure‑as‑code tooling that eliminated manual... 
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    New York, NY
    4 days ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction... 
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    3 days ago
  • $140k - $225k

     ...on private equity, real estate, public debt and equity, infrastructure, life sciences, growth equity, opportunistic, non-investment...  ...on LinkedIn, X, and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability of systems... 
    Full time
    Local area
    Flexible hours

    Blackstone Group

    New York, NY
    4 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview...  ...creating responsible and reliable AI systems, changing...  ...investments in technology infrastructure and world-class talent — along...  ...information available through this site. Capital One Financial... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are...  ...creating responsible and reliable AI systems, changing banking...  ...in technology infrastructure and world-class talent —...  ...information available through this site. Capital One Financial... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands‑on technically while also mentoring a small team of SREs. The InfraSec team collaborates... 
    Local area
    Remote work
    Flexible hours

    MongoDB

    New York, NY
    4 days ago
  • Basis is seeking a Site Reliability Engineer to ensure the reliability, scalability, and performance of our AI-powered accounting platform. You’ll join a high-leverage infrastructure team at the intersection of product and platform, owning systems that keepBasis fast, secure... 

    getbasis.ai

    New York, NY
    15 hours ago
  • $160k - $180k

    Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 

    Socure

    New York, NY
    4 days ago
  •  ...day, members of our team are hands-on scaling-out production infrastructure, building out CI/CD pipelines, and brainstorming with developers...  ...and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS... 
    For contractors

    Varo Money

    New York, NY
    1 day ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    5 days ago
  • $110k - $120k

     ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a...  ...resolve complex production incidents across application and infrastructure layers.Act as the L3 escalation point for operational issues... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    1 day ago
  • $158.5k - $172k

     ...deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you...  ...design, automate, and secure our core infrastructure, deployment platforms, and CI/CD ecosystems...  ...-impact position driving continuous reliability, deep system optimization, and automation... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $167.7k - $245.2k

     ...aspects of the Federal region’s infrastructure and operations, such as...  ....We’re looking for talented engineers with a software or operations...  ...development teams to ensure the reliability, performance and security of...  ...see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    3 days ago
  • $141k - $216.6k

     ...are creating the only platform that combines modern 911 infrastructure with an AI intelligence layer—helping public safety agencies...  ...a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    2 days ago
  • $120k - $200k

     ...DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal...  ...implementation of the company’s global infrastructure architecture. Responsible for cross-region...  ...testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    1 day ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment...  ...and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications... 
    Shift work

    JP Morgan Chase

    New York, NY
    15 hours ago
  • A leading software development company in New York is seeking an entry-level Python engineer to join their team in the Brooklyn office. The role involves working on the AI inference pipeline that powers sophisticated OCR and computer vision products. Candidates should have... 
    Full time
    Work at office

    Mathpix

    New York, NY
    4 days ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible...  ...and systems, architect low latency infrastructure solutions, proactively guard against...  ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Inference Infrastructure. Be the first to apply!