Software Development Engineer, Infrastructure Reliability Engineering
$136.5k - $184.7kAmazon
Join us in building Reflex, the agentic incident-response platform for Amazon's fulfillment network. You'll design and ship production software and AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, generate real-time call intelligence, draft stakeholder communications, and produce structured post-incident records, shifting incident response from a manual, pull-based model to an intelligent, push-based one.Amazon's network of fulfillment centers is the infrastructure that Amazon Robotics runs on. When it degrades, robots stop and packages stop moving. The team manages thousands of high-severity incidents every year, with Incident Managers assembling context across many systems under time pressure before resolution work can even begin. The software you build puts that context in front of responders in seconds and removes the repetitive work that extends incidents today. You'll design the data, evaluation, and feedback mechanisms that make the system measurably better with every incident it touches.This role sits on a software team within Operations Infrastructure Services (OIS), part of Amazon Robotics. You'll stay close to live operations through incident reviews, workflow observation, and call shadowing, and turn what you learn into durable software. Reflex is the immediate focus; as it matures, the same foundations are expected to extend toward shared incident context across organizations, coordinated agent workflows, and carefully guarded automation of recovery validation and repeatable response actions, with every step gated by measurable confidence.Key job responsibilities- Design, build, test, deploy, and operate production services and AI agents on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call intelligence, communications, and post-incident documentation and reporting- Own features end-to-end: from discovery with Incident Managers and resolver teams, through design, implementation, evaluation, deployment, and production operation- Build the foundations that gate agent autonomy: LLM output evaluation, observability and alerting for agents in production, and identity and access controls aligned with Amazon standards- Design the feedback loops through which agents learn: capturing human reviews, corrections, approvals, and incident outcomes as evaluation signal, and turning resolved incidents into structured history that improves recommendations over time- Integrate with the incident lifecycle (ticketing, chat, telemetry, detection feeds, and live call transcription) and model consistent incident state across those systems- Raise the bar on software quality, security, testing, and operational excellence for AI systems acting inside production incident workflowsA day in the lifeYou start by reviewing overnight agent evaluation results and fixing a class of inaccurate triage recommendations, shipping an improvement responders see on the next incident. Later, you pair with an Incident Manager to see how they corrected an agent-drafted call summary, then turn that correction into an automated evaluation case and an agent fix. You might close the day reviewing a design for representing incident state across source systems, or shadowing part of a live bridge call to spot where responders lose time. The operational insight you gather today becomes the software that shortens tomorrow's incident.Amazon offers a full range of benefits to support you and eligible family members, including domestic partners and children. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:1. Medical, Dental, and Vision Coverage2. Maternity and Parental Leave Options3. Paid Time Off (PTO)4. 401(k) PlanIf you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply! About the teamWe're a software engineering team within Operations Infrastructure Services, part of Amazon Robotics. We build incident-response software for major incident management and for the engineering and operations teams that keep Amazon's fulfillment infrastructure healthy. Our users work in high-pressure environments where missing context, unclear ownership, and repetitive manual work directly extend incidents, and we work backward from those problems.Basic qualifications- 3+ years of non-internship professional software development experience- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience- 1+ years of Object Oriented Design experience- Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field- Experience programming with at least one software programming languagePreferred qualification - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniquesAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 136,500.00 - 184,700.00 USD annuallyUSA, VA, Arlington - 143,700.00 - 194,400.00 USD annually
- Senior Site Reliability Engineer - Network Operations (Remote) Fastly 15... ...building and operating the infrastructure that powers the Fastly Edge... .... * Contribute to the development of tools and automation systems... ...to shape roadmaps and software solutions. * Mentor team members...SuggestedRemote jobLocal areaFlexible hoursNight shift
$100k - $150k
...Platform Reliability Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions... ...software engineering principles to infrastructure and operations problems, and continually...SuggestedFull timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...technologies at ByteDance. Build software and tools to improve the reliability and availability of high-speed network infrastructure. Requirements... ...completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline...SuggestedFull time
$143.4k
...non-internship professional software development experience. We require... ...design patterns, reliability, and scaling for new and existing... ...routing and forwarding, traffic engineering, and SDN technologies to... ...services, and terrestrial infrastructure. This role is based in...SuggestedFull timeInternshipRemote workFlexible hours- ...Job Description Are you a Senior Site Reliability Engineer with direct experience operating AI... ...and in a job that is not a traditional infrastructure-only SRE role then this is the... ...application teams to move AI services from development into reliable production environments...SuggestedPermanent employmentRemote work
$10,500 per month
...the world’s leading software for data-driven decisions... ...The Role Software Engineers at Palantir drive... ...AI and world-leading infrastructure that supports mission... ...Designers and Product Reliability Engineers. We also partner... ...with our business development teams (Forward...Work experience placementInternshipWork at officeRemote workWork from homeWorldwide- ...personal and professional development, thereby creating new pathways... .../ applications utilizing infrastructure as code methodologies to solve... ...productivity, security, reliability, and performance Progress... ...years of technical Cloud Engineering experience; U.S. Federal government...Full timeLocal area
$118.28k - $212.87k
...KPMG. Looking ahead, we anticipate continued evolution and success within the practice, fostering both personal and professional development, thereby creating new pathways for growth. In this ever-changing market environment, our professionals must be adaptable and thrive...Full timeH1bLocal area$153.6k - $207.8k
...implementation experience - Strong hands-on development experience with modern programming... ...) - 3+ years of AWS service suite, Infrastructure as Code (Terraform, CloudFormation), Security... ...Architect Professional, DevOps Engineer Professional) preferred - Deep understanding...Flexible hours- ...Responsibilities Own and deliver software projects end-to-end from... ..., Identity, Privacy, Infrastructure, engineering, and product teams on... .... Drive observability, reliability, and operational readiness... ...experience with full-stack development, security, data, privacy,...Full timeFlexible hours
- ...Description The work Software Engineers working on back-end and cloud systems build the... ...platform or DevSecOps teams to balance reliability, security, performance,... ...cloud provider, container orchestration, infrastructure as code, or a particular identity and...Full time
- ...This will enable thousands of engineers at Snowflake and will... ...Utilization) across Snowflake reliably, how do we store it efficiently... ...actively looking for a senior software engineer. If you love... ...developing or using observability infrastructure such as OpenTelemetry,...Full time
$184k - $259.44k
...Software Engineer, Frontier AI Infrastructure Scale AI is seeking a highly skilled and motivated Software Engineer... ...you'd have: Full Stack Development: Proficiency in both front-end... ...Scale, our mission is to develop reliable AI systems for the world's most important...Full timeWork at office3 days per weekEarly shift$115.5k - $184.8k
...Site Reliability Engineer II Join Axon and be a Force for Good. At... ...ecosystem of devices and cloud software. Like our products, we... ...driven configuration, and infrastructure as code that actually reflects... ...support Learning & Development programs Employee Resource...Work at officeRemote work$95k - $171k
...Are you passionate about cutting-edge AI infrastructure? Do you want to build your SRE career... ..., Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless... ...inference platform. As an Site Reliability Engineer II, you will be responsible for:...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$90k - $150k
...honoree, is seeking a SRE Engineer to support our... ...Engineer will support the Infrastructure, Production, and... ...responsible for improving the reliability, availability,... ...ecosystem across development, test, training, and... ...field. ~5 years of software engineering, 3 years...Permanent employmentFull timeContract work$121.4k - $218.6k
...dedicated AI hardware infrastructure. You will be... ...-in-class uptime and reliability of our AI hardware infrastructure... ...density hardware and software infrastructure... ...the earliest stages of development to ensure the reliability... ...Site Reliability Engineer, you will be responsible...Work experience placementWork at office- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable... ..., and collaborate with other teams to ensure seamless infrastructure and application integration. The ideal resource has...Work experience placement
- ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer... ...candidate will bridge the gap between development and operations by applying software engineering principles to infrastructure and operational problems. This role...
$150k - $180k
...seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and... ...the mission- and business-critical infrastructure that powers Umbra's systems. In this... ...~ Expertise in infrastructure and software architecture, capable of designing and...Permanent employmentWork at officeLocal areaRemote workWorldwideFlexible hours- ...a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position... ..., reliable, and secure operation of critical application infrastructure for a federal financial agency. In this operationally...Remote work
$230k - $250k
...Overview GovCIO is hiring a Site Reliability Engineer with an active Secret clearance... ...-critical systems by combining software engineering practices with infrastructure operations expertise. This role... ...strategies. Partner with development teams to improve application...Full timeRemote workFlexible hours- ...Application Solutions Architect /SRE Engineer Important Note : We... ...into our cloud-native development and operations workflows. This... ...expertise in AWS tooling, infrastructure automation, and secure CI/CD... ...to a virtual desktop set up (software) will be provided by Lumen's...For contractorsShift work
$125k - $174.33k
...their experiences working at Coupa. The Impact of a Lead Site Reliability Engineer at Coupa: Joining the team as a Lead Site Reliability... ...automate directory changes, security policy updates, and infrastructure provisioning for our global Active Directory deployment. Lead...- ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027... ..., or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS... ...comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including...
$210k - $230k
...Secret Hybrid schedule IT Infrastructure & Network Engineering & Operations Overview GovCIO... ...currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement... ...candidate will bridge the gap between development and operations, focusing on...Full timeCurrently hiringRemote workFlexible hours$128.5k - $190k
...Senior Site Reliability Engineer Medallia is the pioneer and market leader... ...brings together the infrastructure and applications that power... ...platforms. Partner with software engineering teams to improve... ...using AI-assisted development, automation, or operational...Temporary workWork experience placementLocal area$107k - $220k
...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and... ...rotations to respond to system incidents, collaborate with development teams to improve application reliability and performance,...Full timeContract workTemporary workWork at officeVisa sponsorshipWork visa- ...Site Reliability Engineer Qualifications: ~10+ years of overall... ...including, with hands-on Development and Systems engineering background... ...~ Solid understanding of Software coding techniques and... ...Consultation on Technology infrastructure planning and engineering for...Temporary workImmediate start
- ...Startups 2024,” and Y Combinator’s #1 GovTech startup. About the Role We want a Platform/Infrastructure engineer to help shape how Promise delivers and runs software, and to automate as much of that as possible. The goal is simple, build the systems that enable...Permanent employmentFull timeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Engineer, Infrastructure Reliability Engineering. Be the first to apply!
- software developer positions Arlington, VA
- senior software engineer remote Arlington, VA
- software engineer contract Arlington, VA
- cybersecurity software engineer Arlington, VA
- part time software developer remote Arlington, VA
- junior software developer internship Arlington, VA
- software system engineer Arlington, VA
- software engineer remote Arlington, VA
- software engineer visa sponsorship Arlington, VA
- work from home software developer Arlington, VA





