Software Engineer, Inference
$300kAnthropic
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.
The team has a dual mandate: maximizing compute efficiency to serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms.You may be a good fit if you:
- Have significant software engineering experience, particularly with distributed systems
- Are results-oriented, with a bias towards flexibility and impact
- Pick up slack, even if it goes outside your job description
- Enjoy pair programming (we love to pair!)
- Want to learn more about machine learning systems and infrastructure
- Thrive in environments where technical excellence directly drives both business results and research breakthroughs
- Care about the societal impacts of your work
Strong candidates may also have experience with:
- High-performance, large-scale distributed systems
- Implementing and deploying machine learning systems at scale
- Load balancing, request routing, or traffic management systems
- LLM inference optimization, batching, and caching strategies
- Kubernetes and cloud infrastructure (AWS, GCP)
- Python or Rust
Representative projects:
- Designing intelligent routing algorithms that optimize request distribution across thousands of accelerators
- Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads
- Building production-grade deployment pipelines for releasing new models to millions of users
- Integrating new AI accelerator platforms to maintain our hardware-agnostic competitive advantage
- Contributing to new inference features (e.g., structured sampling, prompt caching)
- Supporting inference for new model architectures
- Analyzing observability data to tune performance based on real-world production workloads
- Managing multi-region deployments and geographic routing for global customers
Deadline to apply: None. Applications will be reviewed on a rolling basis.
The expected base compensation for this position is below. Our total compensation package for full-time employees includes equity, benefits, and may include incentive compensation.
Annual Salary:
$300,000 - $485,000 USD
Logistics
Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
How we're different
We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Come work with us!
Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process
$320k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...the Role Our mandate is to make inference deployment boring and unattended.... ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,...SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference...SuggestedTemporary workFor contractorsWork experience placement$158.1k - $213.8k
...products to Amazon’s customers.We own the complete pipeline for our software, from gathering requirements to development and testing to... ...existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of Object Oriented...SuggestedInternshipWorldwideFlexible hours$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the Role The Cloud Inference team scales and optimizes Claude to serve... ...Fit If You: Have significant software engineering experience, with a strong background...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours$250k - $300k
Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction...SuggestedWork experience placementWork at officeLocal areaImmediate start- ...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft... ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems...Full timeWork experience placementWork at officeLocal areaRemote workHome office
$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI... ...One. ~ Design, develop, test, deploy, and support AI software components including foundation model training, large language...Full timePart timeLocal area$229.9k - $262.4k
...Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large...Full timePart timeLocal area- ...Series A , led by Felicis. About the role We are hiring Software Engineers to join our team. This is an opportunity to join us in-person... ...programs in order to optimize arbitrary user Python code, infers and orchestrates infrastructure implied by the structure of that...Full timeWork at officeFlexible hours
$109k - $145k
...power them. This team enables both internal engineers and customers to monitor, troubleshoot,... ...About the role: As a Software Engineer on the Observability team, you... ...GPU-based systems, large-scale training/inference workloads, or MLOps tooling Why...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...This is a high-ownership generalist role. You'll work across 3 services spanning TypeScript/React, Python backends, and ML/CV inference, all running on Google Cloud and Modal. Comfort moving between product code and ML infrastructure is essential. Physician web portal...Full time
$240k
...defense layer for the AI age and are looking for an exceptional ML engineer to stabilize the system that turns raw signal into decisions —... ...turn them into consistent, trusted decisions. Define how inference works when inputs are incomplete, noisy, or conflicting....Full timeFlexible hours- ...financial ecosystem. Role Description As a Senior Software Engineer , you'll be one of the early technical hires building the systems... ...customer-facing tools Develop infrastructure that serves inference and network analysis results in real time with high accuracy...Full time
$300k - $320k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...: Anthropic is looking for backend software engineers to work across our product... ...our models. You'll partner closely with inference and safeguards to optimize the full...Full timeWork at officeVisa sponsorshipFlexible hours$158.1k - $213.8k
...relevant ad experiences- Partner with engineering, science, and business teams to design,... ...+ years of non-internship professional software development experience- 2+ years of non... ...including transformer architecture, training/inference lifecycles, and optimization...InternshipWorldwideFlexible hours$158.1k - $213.8k
...for building innovation in silicon and software for our AWS customers. We are at the forefront... ...scale with the world’s most talented engineers. Our team covers multiple disciplines... ...chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the...InternshipFlexible hours- ...asynchronous: the result is 10-100× more AI inference per dollar, per watt. We co-design... ...tech investors and built by scientists, engineers, and operators from the labs that built... ...every seniority. The Role As a Software Engineer at Normal, you will build the...Full time
- ...comprehensive LLM serving platform with a focus on end-to-end inference research. You will work with the research lead to pick high‑impact... .... You will collaborate with customers and Forward Deployed Engineers to deploy and tune models, and you will push frontier...
$194k - $239k
...Hover Infrastructure Engineer Hover helps people design, improve, and protect the properties... ...than traditional services do, from LLM inference and agent runtimes to vector stores, GPU... ...or SRE role. Infrastructure here is a software engineering problem. We build systems...Full timeFor contractorsWork at officeLocal areaFlexible hours$150k - $220k
...and enjoy a rewarding career. We are seeking a Senior Software Engineer – Integration to join a new Bruin Platform Modernization... ...with modern IDEs and agentic coding tools (autocomplete, type inference, AI assistants using the SDK as context). ~ Experience building...Contract work$139k - $204k
...training clusters, agent building, and inference at scale, we’re combining forces to serve... ...complex problems at the intersection of software, hardware, and AI, there's never been a... ...effectiveness. About the role As a Software Engineer, you’ll lead efforts to scale our...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...machine learning, a scalable data generation engine, and a partnership track record... ...docking; creating scalable on-demand ML inference infrastructure; developing scalable chemical... ...new features and services Developing software tools to automate or improve processes ranging...Full time
$175k - $250k
...Software Engineer, Machine Learning (MLOps & Data) A Career with Point72’s Surveillance Team On the Knowledge Graph Intelligence... ...full lifecycle of ML models, from data ingestion to production inference, contributing to the design of our next-generation, event-...Full timeWork experience placement$216k - $270k
...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale... ...candidate will have a strong understanding of software engineering principles and practices, as well...Full time$120k - $240k
...is a deep-tech company of scientists and engineers, developing machine learning... ...Own the deployment of custom-tailored software solutions with small, high performing teams... ...learning concepts (eg. model training, model inference, hardware accelerations) What we offer...Full timeWork experience placementWork at officeRemote workFlexible hours- ...We are putting together a tight team of innovative product engineers to join the AI team full stack. This is an in person role in New... ...users ~1+ year of experience on products that deliver AI/ML inference to end users. Can be anywhere in that lifecycle, from training...Full timeContract workWork at office
$196k - $294k
...providers. You will collaborate with a remote and distributed team of engineers to build reliable, low-latency systems that handle rate... ...ensure low-latency responses and stability for high-volume AI inference requests. Collaborate with cross-functional teams,...Full timeRemote workWork from homeFlexible hours$102.5k - $210.6k
...Role Overview: As an Applied AI Engineer III, you will actively engage in your engineering... ...craftsmanship across full-stack software engineering and modern frameworks—together... ...designs and implementations, and owning the inference, token, and cloud cost of what you build...Work at officeLocal areaVisa sponsorshipFlexible hours$137.21k - $185.19k
...Software Engineer Here at Siemens, we take pride in enabling sustainable progress through technology. We do this through empowering customers... ...span the software-hardware boundary, combining low-latency inference pipelines, robust cloud infrastructure, and tightly...Immediate start$150k - $250k
...Forward Deployed Software Engineer New York About Us PhysicsX is a deep-tech company with roots in numerical physics and Formula... ...enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!
- agile software developer New York, NY
- software developer internship no experience New York, NY
- intermediate software engineer New York, NY
- software engineer staff New York, NY
- experienced software developer New York, NY
- software engineer co-op New York, NY
- work from home software developer New York, NY
- software developer no experience New York, NY
- software developer fintech New York, NY
- software data engineer New York, NY



