Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

$300k
Full-time

Anthropic

About Anthropic


Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.



About the role




Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.

The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.

You may be a good fit if you:



  • Have significant software engineering experience, particularly with distributed systems

  • Are results-oriented, with a bias towards flexibility and impact

  • Pick up slack, even if it goes outside your job description

  • Enjoy pair programming (we love to pair!)

  • Want to learn more about machine learning systems and infrastructure

  • Thrive in environments where technical excellence directly drives both business results and research breakthroughs

  • Care about the societal impacts of your work

Strong candidates may also have experience with:



  • High-performance, large-scale distributed systems

  • Implementing and deploying machine learning systems at scale

  • Load balancing, request routing, or traffic management systems

  • LLM inference optimization, batching, and caching strategies

  • Kubernetes and cloud infrastructure (AWS, GCP)

  • Python or Rust

Representative projects:



  • Designing intelligent routing algorithms that optimize request distribution across thousands of accelerators

  • Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads

  • Building production-grade deployment pipelines for releasing new models to millions of users

  • Integrating new AI accelerator platforms to maintain our hardware-agnostic competitive advantage

  • Contributing to new inference features (e.g., structured sampling, prompt caching)

  • Supporting inference for new model architectures

  • Analyzing observability data to tune performance based on real-world production workloads

  • Managing multi-region deployments and geographic routing for global customers

Deadline to apply:  None. Applications will be reviewed on a rolling basis. 

The expected base compensation for this position is below. Our total compensation package for full-time employees includes equity, benefits, and may include incentive compensation.

Annual Salary:

$300,000 - $485,000 USD

Logistics


Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

How we're different


We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.

The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.

Come work with us!


Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about  our policy for using AI in our application process

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in New York, NY vacancy
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...the Role Our mandate is to make inference deployment boring and unattended....  ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    New York, NY
    11 hours ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference... 
    Suggested
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    2 days ago
  • $158.1k - $213.8k

     ...products to Amazon’s customers.We own the complete pipeline for our software, from gathering requirements to development and testing to...  ...existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of Object Oriented... 
    Suggested
    Internship
    Worldwide
    Flexible hours

    Amazon

    New York, NY
    7 hours ago
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the Role The Cloud Inference team scales and optimizes Claude to serve...  ...Fit If You: Have significant software engineering experience, with a strong background... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    1 day ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction... 
    Suggested
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    1 day ago
  •  ...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft...  ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI...  ...One. ~ Design, develop, test, deploy, and support AI software components including foundation model training, large language... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    1 day ago
  •  ...Series A , led by Felicis. About the role We are hiring Software Engineers to join our team. This is an opportunity to join us in-person...  ...programs in order to optimize arbitrary user Python code, infers and orchestrates infrastructure implied by the structure of that... 
    Full time
    Work at office
    Flexible hours

    Chalk

    New York, NY
    11 hours ago
  • $109k - $145k

     ...power them. This team enables both internal engineers and customers to monitor, troubleshoot,...  ...About the role: As a Software Engineer on the Observability team, you...  ...GPU-based systems, large-scale training/inference workloads, or MLOps tooling Why... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Coreweave

    New York, NY
    11 hours ago
  •  ...This is a high-ownership generalist role. You'll work across 3 services spanning TypeScript/React, Python backends, and ML/CV inference, all running on Google Cloud and Modal. Comfort moving between product code and ML infrastructure is essential. Physician web portal... 
    Full time

    Ataraxis AI

    New York, NY
    11 hours ago
  • $240k

     ...defense layer for the AI age and are looking for an exceptional ML engineer to stabilize the system that turns raw signal into decisions —...  ...turn them into consistent, trusted decisions. Define how inference works when inputs are incomplete, noisy, or conflicting.... 
    Full time
    Flexible hours

    Sweep360

    New York, NY
    11 hours ago
  • $300k - $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...: Anthropic is looking for backend software engineers to work across our product...  ...our models. You'll partner closely with inference and safeguards to optimize the full... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    11 hours ago
  •  ...financial ecosystem. Role Description As a Senior Software Engineer , you'll be one of the early technical hires building the systems...  ...customer-facing tools Develop infrastructure that serves inference and network analysis results in real time with high accuracy... 
    Full time

    Cobalt Identity Systems

    New York, NY
    11 hours ago
  •  ...asynchronous: the result is 10-100× more AI inference per dollar, per watt.   We co-design...  ...tech investors and built by scientists, engineers, and operators from the labs that built...  ...every seniority. The Role As a Software Engineer at Normal, you will build the... 
    Full time

    Normal Computing

    New York, NY
    11 hours ago
  • $158.1k - $213.8k

     ...for building innovation in silicon and software for our AWS customers. We are at the forefront...  ...scale with the world’s most talented engineers. Our team covers multiple disciplines...  ...chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the... 
    Internship
    Flexible hours

    Amazon

    New York, NY
    2 days ago
  • $158.1k - $213.8k

     ...relevant ad experiences- Partner with engineering, science, and business teams to design,...  ...+ years of non-internship professional software development experience- 2+ years of non...  ...including transformer architecture, training/inference lifecycles, and optimization... 
    Internship
    Worldwide
    Flexible hours

    Amazon

    New York, NY
    1 day ago
  •  ...comprehensive LLM serving platform with a focus on end-to-end inference research. You will work with the research lead to pick high‑impact...  .... You will collaborate with customers and Forward Deployed Engineers to deploy and tune models, and you will push frontier... 

    modal

    New York, NY
    1 day ago
  • $150k - $220k

     ...and enjoy a rewarding career. We are seeking a Senior Software Engineer – Integration to join a new Bruin Platform Modernization...  ...with modern IDEs and agentic coding tools (autocomplete, type inference, AI assistants using the SDK as context). ~ Experience building... 
    Contract work

    MetTel

    New York, NY
    9 days ago
  • $194k - $239k

     ...Hover Infrastructure Engineer Hover helps people design, improve, and protect the properties...  ...than traditional services do, from LLM inference and agent runtimes to vector stores, GPU...  ...or SRE role. Infrastructure here is a software engineering problem. We build systems... 
    Full time
    For contractors
    Work at office
    Local area
    Flexible hours

    Almaz Capital

    New York, NY
    2 days ago
  • $139k - $204k

     ...training clusters, agent building, and inference at scale, we’re combining forces to serve...  ...complex problems at the intersection of software, hardware, and AI, there's never been a...  ...effectiveness. About the role As a Software Engineer, you’ll lead efforts to scale our... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Weights & Biases

    New York, NY
    2 days ago
  • $175k - $250k

     ...Software Engineer, Machine Learning (MLOps & Data)   A Career with Point72’s Surveillance Team On the Knowledge Graph Intelligence...  ...full lifecycle of ML models, from data ingestion to production inference, contributing to the design of our next-generation, event-... 
    Full time
    Work experience placement

    Point72

    New York, NY
    11 hours ago
  • $216k - $270k

     ...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale...  ...candidate will have a strong understanding of software engineering principles and practices, as well... 
    Full time

    Scale Ai, Inc.

    New York, NY
    11 hours ago
  • $120k - $240k

     ...is a deep-tech company of scientists and engineers, developing machine learning...  ...Own the deployment of custom-tailored software solutions with small, high performing teams...  ...learning concepts (eg. model training, model inference, hardware accelerations) What we offer... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Physicsx

    New York, NY
    11 hours ago
  •  ...machine learning, a scalable data generation engine, and a partnership track record...  ...docking; creating scalable on-demand ML inference infrastructure; developing scalable chemical...  ...new features and services Developing software tools to automate or improve processes ranging... 
    Full time

    Proxima

    New York, NY
    11 hours ago
  •  ...We are putting together a tight team of innovative product engineers to join the AI team full stack. This is an in person role in New...  ...users ~1+ year of experience on products that deliver AI/ML inference to end users. Can be anywhere in that lifecycle, from training... 
    Full time
    Contract work
    Work at office

    Tabs

    New York, NY
    11 hours ago
  • $196k - $294k

     ...providers. You will collaborate with a remote and distributed team of engineers to build reliable, low-latency systems that handle rate...  ...ensure low-latency responses and stability for high-volume AI inference requests. Collaborate with cross-functional teams,... 
    Full time
    Remote work
    Work from home
    Flexible hours

    Vercel

    New York, NY
    11 hours ago
  •  ...an organization can actually use them, not just the data team. We've raised $14M in seed funding, built our team around a strong engineering culture, and are launching publicly after nearly two years in stealth mode. About the Team We're building an agentic AI platform... 
    Full time
    Immediate start
    Home office
    Flexible hours

    Merciv

    New York, NY
    11 hours ago
  • $160k - $220k

     ...around the world, Via is recognized as the leading transportation technology and service provider globally. As a Full-Stack Software Engineer on the Remix engineering team, you’ll build software cities rely on to design and improve public transportation systems with... 
    Full time
    2 days per week
    3 days per week

    Via

    New York, NY
    11 hours ago
  •  ...from first request to production deployment quickly and confidently. About the Role We’re looking for full stack and frontend engineers to help define and build the next generation of OpenAI’s developer experience. In this role, you’ll work across frontend product... 
    Full time
    Internship

    OpenAI

    New York, NY
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!