Sr. Applied Scientist, Infrastructure Reliability
$158.8k - $214.8kAmazon Locker
Are you driven by innovation and complex problem-solving? At Infrastructure Reliability, we build scalable solutions that ensure the reliability of Amazon's critical systems. Our team develops and operates tools that detect and prevent outages to maintain high availability across Amazon's global infrastructure. Join us to architect solutions that directly impact millions of customers, with the resources and support to make meaningful contributions.The team at Amazon is responsible for building intelligent and real-time insights into service-to-service communications, network traffic, and event correlation across hundreds of Amazon's critical fulfillment and robotics services. Our solutions support visibility into anomalous service behavior to prevent and quickly recover from incidents, ensuring high availability to keep the Customer Promise.We are seeking a talented Senior Applied Scientist to invent the next generation of agentic observability solutions at Amazon scale. In this role, you will define, lead, and build the science behind intelligent systems that reason about complex network and infrastructure telemetry, autonomously detect anomalies, and drive automated remediation. You will own the scientific direction end to end, partnering closely with engineering, product, and Network Development Engineers to translate a long-term science vision into concrete research and delivery roadmaps.Working backwards from the needs of our customers and operations teams, you will take the lead on ambiguous, high-impact problems where neither the problem nor the solution is well defined, and deliver production systems that improve infrastructure availability at scale. You will invent new methods, drive their adoption across multiple teams, and remain deeply hands-on with the hardest technical challenges.Key job responsibilities* Define and own the science vision for agentic observability, translating it into research and engineering roadmaps in partnership with product and engineering leaders.* Build ML and agentic AI systems that autonomously detect, classify, and correlate infrastructure anomalies across network, compute, and service layers at Amazon scale.* Design and develop models for event correlation, root cause analysis, and predictive failure detection using time-series analysis, graph-based methods, and deep learning.* Own the agentic architecture for automated observability workflows, including planning, tool integration, long-horizon reasoning, and multi-agent orchestration.* Define and curate the datasets and evaluation methodologies needed to train, benchmark, and continuously improve detection and classification systems.* Stay deeply hands-on: write production-quality, critical-path code and build core components that take systems from prototype to launch.Partner with Network Development Engineers and operations teams to ground science solutions in real-world infrastructure behavior and operational needs.* Mentor scientists and engineers, raise the science bar, and represent the team in the internal and external scientific community through publications and presentations.A day in the lifeYou will solve real-world problems by analyzing large-scale network telemetry and operational data, designing experiments and simulations, and developing ML models that detect and prevent infrastructure incidents. Your work requires close collaboration with engineers, Network Development Engineers, product managers, and operations leaders across the organization. You will prepare written and verbal presentations to share insights with audiences of varying technical sophistication, and you will iterate rapidly between research and production deployment.Basic qualifications- 3+ years of building machine learning models for business application experience- PhD, or Master's degree and 6+ years of applied research experience- Experience programming in Java, C++, Python or related language- Experience with neural deep learning methods and machine learningPreferred qualification - Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.- Experience in several of the following areas: machine learning, statistics, deep learning, natural language processing, or information retrieval- Experience with time-series analysis, anomaly detection, or graph-based ML methods- Experience with agentic AI architectures, LLMs, or multi-agent systems- Experience applying ML to infrastructure, networking, or observability domainsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 158,800.00 - 214,800.00 USD annually
$114.6k - $234.6k
..., and responsible AI. As a Senior Applied Scientist on the team, you will independently own... ...evaluation protocols and reusable infrastructure. You will examine more than aggregate... ...contamination, robustness, cost, latency, reliability, safety, and operational constraints....SeniorTemporary workFlexible hours$169.8k - $355.4k
...recognition. Our goal is to enable customers to apply computer vision to solve their... ...opportunity to work with teams of applied scientists and engineers to deliver high quality... ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...SeniorTemporary workFlexible hours$126.2k - $264.1k
...production-ready AI solutions. Leads end-to-end applied research across next-generation facial... ...for real-world deployment. Writes reliable, well-documented code and collaborates... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power...SuggestedTemporary workFlexible hoursShift work$169.3k - $304.7k
...generation of the network? Join our Network Infrastructure SRE team! Our team designs, develops... ...fast, efficient, scalable, and reliable routing software and infrastructure that... ...financial wellness; Eligibility requirements apply. Equal Employment Opportunity Rights...SuggestedWork experience placementWork at office$142k - $194k
...scalable distributed systems or cloud infrastructure. Strong programming skills in Java,... ...virtualization, storage, networking, security, reliability, or platform automation. You should... ...and partition-handling policies. We apply load shedding, throttling, and rate...SuggestedFull time$146.3k - $306.4k
...Establishes KPIs and advanced telemetry; applies formal verification for complex... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... ...needs within the business unit.System Reliability Performance:Define key performance...Temporary workFlexible hours- ...through six business verticals: Testing, Inspection & Consulting; Infrastructure; Utility Services; Environmental Health Sciences; Buildings &... ...location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not...SeniorFull timeWork at officeLocal areaLong distanceNight shiftWeekend work
$114.6k - $234.6k
...elasticity, durability, availability, reliability, and operational-readiness requirements... ...operational excellence mentoring. Build Infrastructure as Code and operational automation for... ..., rollbacks, and change management. Apply encryption, access controls, security...Full timeWork at officeRelocation packageFlexible hours$114.6k - $234.6k
...large-scale distributed systems or cloud infrastructure. We need strong experience designing... ..., performance and load testing, reliability engineering, and production incident response... ..., and recovery-oriented design. We apply distributed-systems tradeoffs for network...Full timeWork at officeWorldwideRelocationRelocation packageFlexible hours- ...Infrastructure Architect Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates... ...issues spanning multiple services, applying advanced investigation and debugging techniques...Temporary workImmediate startFlexible hoursShift work
$114.6k - $234.6k
Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises... ...set the technical direction for their reliability, correctness, security, and operational... ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for...Temporary workWork at officeWorldwideRelocationRelocation packageFlexible hours$188.1k - $254.5k
Amazon's Reliability and Maintenance Engineering (RME) Workforce Excellence team seeks an experienced... ...RME Workforce Excellence team today!The Sr. Manager, Biz Ops will develop and... ...information. If the country/region you’re applying in isn’t listed, please contact your Recruiting...SeniorFlexible hours$114.6k - $234.6k
...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting... ...‑management plans.We are seeking a Core Infrastructure Engineer to design, build, and operate the distributed systems that power reliable, secure, and scalable cloud services....Temporary workFlexible hours$114.6k - $234.6k
...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... ....System Design & Architecture - System Reliability Design:-Build and design fault-tolerant...Temporary workFlexible hoursShift work$121.4k - $218.6k
...that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... ...wellness; Eligibility requirements apply. Equal Employment Opportunity Rights...SeniorWork experience placementWork at office$142.3k - $195.7k
...storage architecture and long-term data infrastructure strategy. This role establishes... ...self-service capabilities that improve reliability and reduce operational effort. Partner... ...valid job requirements. This policy shall apply to all employment actions, including but...Permanent employmentFull timeTemporary workApprenticeshipWork at officeRemote workWork from homeHome office$142k - $194k
...want 3-5+ years of software engineering experience building reliable services, infrastructure, or distributed systems. We need strong programming... ...using redundancy, replication, and automatic failover. Apply recovery-oriented principles and implement retries,...SeniorFull time$123k - $195k
Job Title Senior Cloud Infrastructure Engineer Job Description You will help maintain and secure... ...Cloud High services, helping ensure reliable, secure, and compliant cloud operations.... ...Interested candidates are encouraged to apply as soon as possible to ensure consideration...SeniorFull timeWork at officeImmediate startWork visaRelocation package3 days per week- ...involved in executing large-scale, complex infrastructure projects? Connico is a national... ...through clear communication, sound judgment, reliability, and expert guidance. Participate in... ...you will also have the opportunity to apply your expertise across an expanding range...SeniorFull timeFor contractorsWork at officeRemote work
$79.2k - $209.5k
...systems, fault tolerance, algorithms, system reliability, and real-time state management. Experience applying machine learning or AI frameworks and techniques... ...systems that correlate live network and infrastructure signals, detect customer impact, and initiate...SeniorFull timeFlexible hours$79.2k - $209.5k
...customer engineering teams and Oracle’s cloud infrastructure, networking, engineering, operations, and support organizations to ensure reliable Day 2 operations, resolve critical issues... ...expansion of GPU environments.You will apply strong technical judgment to coordinate...SeniorTemporary workFlexible hours$81.1k - $187k
Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity... ...concepts.Patching and Software MaintenanceExperience applying operating system, middleware, and application patches...SeniorTemporary workFlexible hours$79.2k - $209.5k
...redundancy, replication, automatic failover), applies recovery‑oriented principles, and... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... ....System Design & Architecture - System Reliability Design:-Collaborates with team to build...SeniorTemporary workFlexible hours$146.3k - $306.4k
...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... ....System Design & Architecture - System Reliability Design:Manages the strategy for... ...expertise in new areas, coaching them to apply learnings to advance the organization....SeniorTemporary workFlexible hours- ...energy permitting, with additional opportunities to contribute to onshore wind, electrical transmission, oil & gas, and technology infrastructure projects. As a bat biologist, you will manage bat related tasks including development of proposal scope and budgets, task...SeniorRemote workNight shift
$114.6k - $234.6k
...building and operating large-scale, highly distributed service infrastructure. You should have experience in an operational environment... ...simplicity and resilience in mind. We improve service reliability, latency, and operational automation through intelligent tooling...Full timeWork at officeRelocationFlexible hours$79.2k - $209.5k
...validate, and monitor liquid-cooled GPU infrastructure across factory environments. Develops secure... ...data-center ingestion, and accelerate reliable hyperscale AI infrastructure delivery.... ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for...SeniorTemporary workFlexible hours$146.3k - $306.4k
Oracle Cloud Infrastructure (OCI) is seeking a Lead Principal Core Infrastructure Engineer to join our Networking organization, focusing... ...architect, build, and enhance solutions to safeguard the reliability and availability of OCI’s global cloud services. The ideal candidate...Temporary workFlexible hours$19 per hour
...an environmental testing laboratory as a Scientist 1, conducting hands-on analysis of air,... ...supporting the production of accurate, reliable scientific data. This role is suited to... ...environmental testing. Ability to learn and apply established analytical methods and...Hourly payFull timeFlexible hours$101.55 - $156.73 per hour
...Systems (Varian Aria). Conversant in network topology, server infrastructure, backup system design. General programming experience (SQL,... ...engaged. Learn more about our comprehensive benefits package By applying for a position with Intermountain, I acknowledge that I will...Hourly payLocal areaRemote workMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Applied Scientist, Infrastructure Reliability. Be the first to apply!
- safety scientist Nashville, TN
- graduate scientist Nashville, TN
- research scientist - biology Nashville, TN
- scientist 1 Nashville, TN
- research scientist Nashville, TN
- scientist ii Nashville, TN
- manufacturing scientist Nashville, TN
- applied scientist Nashville, TN
- operations research scientist Nashville, TN
- application scientist Nashville, TN


