Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Engineer, Infrastructure Reliability Engineering

$136.5k - $184.7k

Amazon

Join us in building Reflex, the agentic incident-response platform for Amazon's fulfillment network. You'll design and ship production software and AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, generate real-time call intelligence, draft stakeholder communications, and produce structured post-incident records, shifting incident response from a manual, pull-based model to an intelligent, push-based one.Amazon's network of fulfillment centers is the infrastructure that Amazon Robotics runs on. When it degrades, robots stop and packages stop moving. The team manages thousands of high-severity incidents every year, with Incident Managers assembling context across many systems under time pressure before resolution work can even begin. The software you build puts that context in front of responders in seconds and removes the repetitive work that extends incidents today. You'll design the data, evaluation, and feedback mechanisms that make the system measurably better with every incident it touches.This role sits on a software team within Operations Infrastructure Services (OIS), part of Amazon Robotics. You'll stay close to live operations through incident reviews, workflow observation, and call shadowing, and turn what you learn into durable software. Reflex is the immediate focus; as it matures, the same foundations are expected to extend toward shared incident context across organizations, coordinated agent workflows, and carefully guarded automation of recovery validation and repeatable response actions, with every step gated by measurable confidence.Key job responsibilities- Design, build, test, deploy, and operate production services and AI agents on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call intelligence, communications, and post-incident documentation and reporting- Own features end-to-end: from discovery with Incident Managers and resolver teams, through design, implementation, evaluation, deployment, and production operation- Build the foundations that gate agent autonomy: LLM output evaluation, observability and alerting for agents in production, and identity and access controls aligned with Amazon standards- Design the feedback loops through which agents learn: capturing human reviews, corrections, approvals, and incident outcomes as evaluation signal, and turning resolved incidents into structured history that improves recommendations over time- Integrate with the incident lifecycle (ticketing, chat, telemetry, detection feeds, and live call transcription) and model consistent incident state across those systems- Raise the bar on software quality, security, testing, and operational excellence for AI systems acting inside production incident workflowsA day in the lifeYou start by reviewing overnight agent evaluation results and fixing a class of inaccurate triage recommendations, shipping an improvement responders see on the next incident. Later, you pair with an Incident Manager to see how they corrected an agent-drafted call summary, then turn that correction into an automated evaluation case and an agent fix. You might close the day reviewing a design for representing incident state across source systems, or shadowing part of a live bridge call to spot where responders lose time. The operational insight you gather today becomes the software that shortens tomorrow's incident.Amazon offers a full range of benefits to support you and eligible family members, including domestic partners and children. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:1. Medical, Dental, and Vision Coverage2. Maternity and Parental Leave Options3. Paid Time Off (PTO)4. 401(k) PlanIf you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply! About the teamWe're a software engineering team within Operations Infrastructure Services, part of Amazon Robotics. We build incident-response software for major incident management and for the engineering and operations teams that keep Amazon's fulfillment infrastructure healthy. Our users work in high-pressure environments where missing context, unclear ownership, and repetitive manual work directly extend incidents, and we work backward from those problems.Basic qualifications- 3+ years of non-internship professional software development experience- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 1+ years of software development engineer or related occupational experience- 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience- 1+ years of Object Oriented Design experience- Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field- Experience programming with at least one software programming languagePreferred qualification - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniquesAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 136,500.00 - 184,700.00 USD annuallyUSA, VA, Arlington - 143,700.00 - 194,400.00 USD annually

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Software Development Engineer, Infrastructure Reliability Engineering in Arlington, VA vacancy
  • Senior Site Reliability Engineer - Network Operations (Remote) Fastly 15...  ...building and operating the infrastructure that powers the Fastly Edge...  .... * Contribute to the development of tools and automation systems...  ...to shape roadmaps and software solutions. * Mentor team members... 
    Suggested
    Remote job
    Local area
    Flexible hours
    Night shift

    Fastly

    Mc Lean, VA
    5 days ago
  • $100k - $150k

     ...Platform Reliability Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions...  ...software engineering principles to infrastructure and operations problems, and continually... 
    Suggested
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Washington DC
    a month ago
  •  ...technologies at ByteDance. Build software and tools to improve the reliability and availability of high-speed network infrastructure. Requirements...  ...completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline... 
    Suggested
    Full time

    ByteDance

    Washington DC
    9 days ago
  • $143.4k

     ...non-internship professional software development experience. We require...  ...design patterns, reliability, and scaling for new and existing...  ...routing and forwarding, traffic engineering, and SDN technologies to...  ...services, and terrestrial infrastructure. This role is based in... 
    Suggested
    Full time
    Internship
    Remote work
    Flexible hours

    Amazon Kuiper Manufacturing Enterprises LLC

    Washington DC
    8 days ago
  •  ...Job Description Are you a Senior Site Reliability Engineer with direct experience operating AI...  ...and in a job that is not a traditional infrastructure-only SRE role then this is the...  ...application teams to move AI services from development into reliable production environments... 
    Suggested
    Permanent employment
    Remote work

    HTC Global Services Inc

    Washington DC
    2 days ago
  • $10,500 per month

     ...the world’s leading software for data-driven decisions...  ...The Role Software Engineers at Palantir drive...  ...AI and world-leading infrastructure that supports mission...  ...Designers and Product Reliability Engineers. We also partner...  ...with our business development teams (Forward... 
    Work experience placement
    Internship
    Work at office
    Remote work
    Work from home
    Worldwide

    Palantir Technologies

    Washington DC
    1 day ago
  •  ...personal and professional development, thereby creating new pathways...  .../ applications utilizing infrastructure as code methodologies to solve...  ...productivity, security, reliability, and performance Progress...  ...years of technical Cloud Engineering experience; U.S. Federal government... 
    Full time
    Local area

    KPMG

    Washington DC
    1 day ago
  • $118.28k - $212.87k

     ...KPMG. Looking ahead, we anticipate continued evolution and success within the practice, fostering both personal and professional development, thereby creating new pathways for growth. In this ever-changing market environment, our professionals must be adaptable and thrive... 
    Full time
    H1b
    Local area

    KPMG

    Washington DC
    8 days ago
  • $153.6k - $207.8k

     ...implementation experience - Strong hands-on development experience with modern programming...  ...) - 3+ years of AWS service suite, Infrastructure as Code (Terraform, CloudFormation), Security...  ...Architect Professional, DevOps Engineer Professional) preferred - Deep understanding... 
    Flexible hours

    Amazon Web Services, Inc.

    Arlington, VA
    12 hours ago
  •  ...Responsibilities Own and deliver software projects end-to-end from...  ..., Identity, Privacy, Infrastructure, engineering, and product teams on...  .... Drive observability, reliability, and operational readiness...  ...experience with full-stack development, security, data, privacy,... 
    Full time
    Flexible hours

    Twitch

    Washington DC
    27 days ago
  •  ...Description The work Software Engineers working on back-end and cloud systems build the...  ...platform or DevSecOps teams to balance reliability, security, performance,...  ...cloud provider, container orchestration, infrastructure as code, or a particular identity and... 
    Full time

    Opendatajobs

    Washington DC
    1 day ago
  •  ...This will enable thousands of engineers at Snowflake and will...  ...Utilization) across Snowflake reliably, how do we store it efficiently...  ...actively looking for a senior software engineer. If you love...  ...developing or using observability infrastructure such as OpenTelemetry,... 
    Full time

    Snowflake

    Washington DC
    1 day ago
  • $184k - $259.44k

     ...Software Engineer, Frontier AI Infrastructure Scale AI is seeking a highly skilled and motivated Software Engineer...  ...you'd have: Full Stack Development: Proficiency in both front-end...  ...Scale, our mission is to develop reliable AI systems for the world's most important... 
    Full time
    Work at office
    3 days per week
    Early shift

    Scale AI

    Washington DC
    3 days ago
  • $115.5k - $184.8k

     ...Site Reliability Engineer II Join Axon and be a Force for Good. At...  ...ecosystem of devices and cloud software. Like our products, we...  ...driven configuration, and infrastructure as code that actually reflects...  ...support Learning & Development programs Employee Resource... 
    Work at office
    Remote work

    Axon

    Washington DC
    40 minutes ago
  • $95k - $171k

     ...Are you passionate about cutting-edge AI infrastructure? Do you want to build your SRE career...  ..., Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless...  ...inference platform. As an Site Reliability Engineer II, you will be responsible for:... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    3 days ago
  • $90k - $150k

     ...honoree, is seeking a SRE Engineer to support our...  ...Engineer will support the Infrastructure, Production, and...  ...responsible for improving the reliability, availability,...  ...ecosystem across development, test, training, and...  ...field. ~5 years of software engineering, 3 years... 
    Permanent employment
    Full time
    Contract work

    Spatial Front

    Arlington, VA
    12 hours ago
  • $121.4k - $218.6k

     ...dedicated AI hardware infrastructure. You will be...  ...-in-class uptime and reliability of our AI hardware infrastructure...  ...density hardware and software infrastructure...  ...the earliest stages of development to ensure the reliability...  ...Site Reliability Engineer, you will be responsible... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable...  ..., and collaborate with other teams to ensure seamless infrastructure and application integration. The ideal resource has... 
    Work experience placement

    Samprasoft

    Washington DC
    12 hours ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer...  ...candidate will bridge the gap between development and operations by applying software engineering principles to infrastructure and operational problems. This role... 

    Software Technology Inc

    Washington DC
    4 days ago
  • $150k - $180k

     ...seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and...  ...the mission- and business-critical infrastructure that powers Umbra's systems. In this...  ...~ Expertise in infrastructure and software architecture, capable of designing and... 
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    1 hour ago
  •  ...a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position...  ..., reliable, and secure operation of critical application infrastructure for a federal financial agency. In this operationally... 
    Remote work

    Elevate Government Solutions

    Washington DC
    1 day ago
  • $230k - $250k

     ...Overview GovCIO is hiring a Site Reliability Engineer with an active Secret clearance...  ...-critical systems by combining software engineering practices with infrastructure operations expertise. This role...  ...strategies. Partner with development teams to improve application... 
    Full time
    Remote work
    Flexible hours

    Govcio

    Arlington, VA
    3 days ago
  •  ...Application Solutions Architect /SRE Engineer Important Note : We...  ...into our cloud-native development and operations workflows. This...  ...expertise in AWS tooling, infrastructure automation, and secure CI/CD...  ...to a virtual desktop set up (software) will be provided by Lumen's... 
    For contractors
    Shift work

    Lumen Solutions Group, Inc.

    Washington DC
    4 days ago
  • $125k - $174.33k

     ...their experiences working at Coupa. The Impact of a Lead Site Reliability Engineer at Coupa: Joining the team as a Lead Site Reliability...  ...automate directory changes, security policy updates, and infrastructure provisioning for our global Active Directory deployment. Lead... 

    GrabJobs

    Arlington, VA
    4 days ago
  •  ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027...  ..., or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS...  ...comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including... 

    Georgia IT Inc

    Washington DC
    4 days ago
  • $210k - $230k

     ...Secret Hybrid schedule IT Infrastructure & Network Engineering & Operations Overview GovCIO...  ...currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement...  ...candidate will bridge the gap between development and operations, focusing on... 
    Full time
    Currently hiring
    Remote work
    Flexible hours

    Govcio

    Arlington, VA
    3 days ago
  • $128.5k - $190k

     ...Senior Site Reliability Engineer Medallia is the pioneer and market leader...  ...brings together the infrastructure and applications that power...  ...platforms. Partner with software engineering teams to improve...  ...using AI-assisted development, automation, or operational... 
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    1 day ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and...  ...rotations to respond to system incidents, collaborate with development teams to improve application reliability and performance,... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall...  ...including, with hands-on Development and Systems engineering background...  ...~ Solid understanding of Software coding techniques and...  ...Consultation on Technology infrastructure planning and engineering for... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    12 hours ago
  •  ...Startups 2024,” and Y Combinator’s #1 GovTech startup. About the Role We want a Platform/Infrastructure engineer to help shape how Promise delivers and runs software, and to automate as much of that as possible. The goal is simple, build the systems that enable... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Promise

    Washington DC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Engineer, Infrastructure Reliability Engineering. Be the first to apply!