Software Engineer, AI Reliability
$325kAnthropic
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the Role
AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects.
Reliability here is an emergent phenomenon that transcends any single team's boundaries, so someone has to zoom out and look at the whole picture. That's us -- and it means few teams at Anthropic offer this kind of dynamic, cross-cutting exposure to the systems that matter most.
Claude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it.
Responsibilities:
Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
Design and implement monitoring and observability systems across the token path.
Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers
Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
Support the reliability of safeguard model serving -- critical for both site reliability and Anthropic's safety commitments.
You may be a good fit if you:
Have strong distributed systems, infrastructure, or reliability backgrounds -- we're looking for reliability-minded software engineers and SREs.
Are curious and brave -- comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don't have deep expertise yet.
Think holistically about how systems compose and where the seams are.
Can build lasting relationships across teams -- our engagement model depends on being welcomed as teammates, not outsiders with opinions.
Care about users and feel ownership over outcomes, even for systems you don't own.
Have excellent communication and collaboration skills -- you'll be partnering across the entire company.
Bring diverse experience -- the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between.
Strong candidates may also:
Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems
Have experience operating large-scale model serving or training infrastructure (>1000 GPUs).
Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium).
Understand ML-specific networking optimizations like RDMA and InfiniBand.
Have expertise in AI-specific observability tools and frameworks.
Have experience with chaos engineering and systematic resilience testing.
Have contributed to open-source infrastructure or ML tooling.
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Annual Salary:
$325,000 - $485,000 USD
Logistics
Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit directly for confirmed position openings.How we're different
We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Come work with us!
Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process
$200k - $300k
Hudson River Trading (HRT) is seeking a Software Engineer focused on GPU reliability to join our Systems Development team. The Systems Development team builds... ...we’d love to get to know you.Please be advised: Use of AI tools during interviews or assessments is strictly...SuggestedWork at officeLocal areaImmediate start$140k - $225k
...and Instagram.Role: Blackstone's Site Reliability Engineering team is responsible for improving the reliability... ...and engineers directly and through AI assistantsImplement instrumentation and... ...either, Infrastructure Engineering, Software Engineering, DevOps Engineering or...SuggestedFull timeLocal areaFlexible hours$140k - $215k
...with the world’s most advanced AI-native platform. We work on... ...our Core Platform and Embedded Reliability charters: building the... ...embedding directly with product engineering teams and their leadership to... ...source community; evangelize software engineering best practices, especially...SuggestedFull timeWork experience placementWork at officeLocal area2 days per week3 days per week$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively... ...US Deloitte Technology Product Engineering has modernized software and product delivery, creating a scalable, cost-effective...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...unique ideas and constant iteration. As AI makes execution faster and tactics... ...our Glassdoor page! The Role As a Software Engineer on one of our Infrastructure or Platform... ...like. What You'll Work On Scale, reliability, and performance. Attack DB contention...SuggestedFull time
- ...teams function in the age of AI. We believe AI isn’t just the... ...About the Role Growth engineering builds the systems that power... ...What You’ll Do Build software to grow Ramp to its next millions... ...Experience shipping fast, reliable, real-time applications using...Remote jobFull timeWork from homeHome officeRelocation packageFlexible hours
- ...Ernest Shackleton’s. No formal experience engineering specifically in crypto, traditional... ...what we call N of 1 products (like onchain AI data analysts!) We start by powering the... ...quickly You’re a self-starter who values reliability, clarity, and impact over hierarchy...Full time
$140k - $260k
...mission to help companies understand and control their AI presence. As a Backend Engineer, you will build and optimize the infrastructure that powers... ...data Ensure system security, performance, and reliability at every level of the stack Work closely with front-...Full timeWork at officeVisa sponsorship- ...Description As a engineer at Bask you will work directly with the... ...and CTO to build and design the software infrastructure that will serve... ...Engineering at Bask Health is AI-first and builder-led. Work... ...operational behavior feel simple and reliable. Work AI-first with Cursor...Full time
$182.75k - $278.5k
...states run their practice on our software, serving over 1 million... ...end ownership. You’re building reliable abstraction layers over a fragmented... ...could work on: The Modern AI-enabled Clinical System:... ...improve your team's engineering velocity - whether that's refining...Full timeFlexible hoursShift work- ...As a Full Stack Software Engineer at Regard, you’ll be involved in all stages of the product development... .... By delivering innovative and reliable software, you’ll help drive the success... ...class healthcare to everyone. Regard is an AI-powered Proactive Documentation platform...Full timeWork at officeLocal areaHome officeVisa sponsorshipRelocation package
- ...bring real time economic optimization and AI prediction to every energy and... ...decisions every minute that determine cost, reliability, and margin, but the signals that matter... ...Role Overview As a Senior Full Stack Software Engineer, you will play a key role in designing,...Full timeLive inWork at officeRemote workVisa sponsorshipFlexible hours
- ...industries are falling behind on AI adoption. Their workflows can’... ...models to make their outputs reliable, traceable, and verifiable.... ...platform and built the analytics engine behind $100M+ contracts. Our... ...that moves fast but ships broken software, this is different. If you've...Full timeWork at office
$160k - $250k
...Overview As an early full-stack engineer, you’ll help build Conductor’s... ...Build polished, reliable user-facing features in a Next... ...~3+ years as a professional software engineer ~ Strong communication... ...~ Experience with LLMs and AI engineering techniques Nice...Full timeContract workLocal areaRemote work$105k - $115k
...bring real time economic optimization and AI prediction to every energy and... ...decisions every minute that determine cost, reliability, and margin, but the signals that matter... ...environments. Role Overview As a Full Stack Software Engineer, you will improve and expand CVector's...Full timeLive inWork at officeVisa sponsorshipFlexible hours- A cutting-edge AI company located in New York is looking for a Platform Engineer to build and operate core infrastructure. This role involves defining best practices for cloud architecture, ensuring reliability and scalability. Ideal candidates will have experience with...
- ...experiences that make OpenAI’s capabilities reliable, understandable, and useful in production... ...’re looking for full stack and frontend engineers to help define and build the next... ...and developer-facing teams to make complex AI capabilities simple to understand, easy to...Full timeInternship
- ...compliance for your team in just 10 mins. Using AI Agents, we automate all state tax... ...next phase, we have some amazing infra, engineering, product, and GTM opportunities ahead of... ...payroll infrastructure that is fast and reliable. You will be focused on ensuring that our...Full timeLive inWork at officeRemote work
- ...unique ideas and constant iteration. As AI makes execution faster and tactics easier... ...: We're looking for a backend product engineer to own the systems behind the parts of Clay... ...use it, and then making it faster, more reliable, and more capable. You'll wear both a PM...Full time
$155k - $175k
...Role Overview The Senior Full Stack Software Engineer will report directly to the Director... ...deliver business value quickly and ensure reliability for the patients and clinicians that... ...We may use artificial intelligence (AI) tools to support parts of the hiring process...Remote jobFull time$300k - $320k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...growing group of committed researchers, engineers, policy experts, and business... ...Anthropic is looking for backend software engineers to work across our product...Full timeWork at officeVisa sponsorshipFlexible hours- ...be part of the core team building truly AI-native helpful experiences across the... ...looking for a stellar & highly ambitious Software Engineer specializing in Go as a core employee... ...existing backend systems for performance and reliability Write clean, maintainable, and well-...Full timeVisa sponsorshipFlexible hours
$280k
...Job description Full Stack Software Engineer — Our Client (Brooklyn, NY | On-Site | Visa Sponsorship... ...) Industry: Software Development · AI · Data Engineering About the Role... ...technical initiatives and deliver scalable, reliable solutions Collaborate closely with...Contract workLocal areaRelocationVisa sponsorshipRelocation package$160k - $190k
...0x better experiences, especially in the AI age. We’re building the next generation of... ...and more. We’re hiring a world-class Engineer Pinwheel is looking for a creative product... ...high product velocity with quality and reliability Contribute to iterative technical...Full timeTemporary workWork at officeFlexible hours3 days per week$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure.... ...energy with the resources of a large software company.What you'll doThis is a role for...Full timeTemporary workLocal areaRemote workWorldwide$141k - $216.6k
...with our ecosystem of devices and cloud software. Like our products, we work better... ...combines modern 911 infrastructure with an AI intelligence layer—helping public safety... ...world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability...Work experience placementWork at office$167.7k - $245.2k
...team is defining the future of AI resilience. Our team provides... ...as intended, improving reliability and reducing risks. This unified... ...As a Senior Site Reliability Engineer (SRE), you will build, operate... ...Code tools• Collaborate with software engineers and customers to design...Full timeTemporary workLocal areaFlexible hours2 days per week$138.1k - $198.2k
...that simply works. The SRE Engineering Enablement Team supports our... ...work. We support the entire software development lifecycle (SDLC),... ...Cisco. Your Impact As a Site Reliability Engineer, you will be at the... ...protect organizations in the AI era - and beyond. We’ve been...Permanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$167.7k - $245.2k
...even the ones they don’t own. Powered by AI and an unmatched set of cloud,... ...effective.We’re looking for talented engineers with a software or operations background, experienced... ...application development teams to ensure the reliability, performance and security of our infrastructure...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$123k - $165k
...Department/Group OverviewOur engineering fleet is a horizontal set of... .... Our specific team provides reliability engineering and operational support... ...workflows.Collaborate with software engineering teams to... ...teams.Experience using modern AI-assisted development tools (e...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, AI Reliability. Be the first to apply!
- senior robotics software engineer New York, NY
- software system engineer New York, NY
- part time software developer New York, NY
- fall software engineering internship New York, NY
- security software engineer New York, NY
- intel software engineer New York, NY
- software developer fintech New York, NY
- new graduate software engineer New York, NY
- software development engineer aws New York, NY
- information technology software engineer New York, NY

