Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Applied Scientist, Infrastructure Reliability

$158.8k - $214.8k

Amazon Locker

Are you driven by innovation and complex problem-solving? At Infrastructure Reliability, we build scalable solutions that ensure the reliability of Amazon's critical systems. Our team develops and operates tools that detect and prevent outages to maintain high availability across Amazon's global infrastructure. Join us to architect solutions that directly impact millions of customers, with the resources and support to make meaningful contributions.The team at Amazon is responsible for building intelligent and real-time insights into service-to-service communications, network traffic, and event correlation across hundreds of Amazon's critical fulfillment and robotics services. Our solutions support visibility into anomalous service behavior to prevent and quickly recover from incidents, ensuring high availability to keep the Customer Promise.We are seeking a talented Senior Applied Scientist to invent the next generation of agentic observability solutions at Amazon scale. In this role, you will define, lead, and build the science behind intelligent systems that reason about complex network and infrastructure telemetry, autonomously detect anomalies, and drive automated remediation. You will own the scientific direction end to end, partnering closely with engineering, product, and Network Development Engineers to translate a long-term science vision into concrete research and delivery roadmaps.Working backwards from the needs of our customers and operations teams, you will take the lead on ambiguous, high-impact problems where neither the problem nor the solution is well defined, and deliver production systems that improve infrastructure availability at scale. You will invent new methods, drive their adoption across multiple teams, and remain deeply hands-on with the hardest technical challenges.Key job responsibilities* Define and own the science vision for agentic observability, translating it into research and engineering roadmaps in partnership with product and engineering leaders.* Build ML and agentic AI systems that autonomously detect, classify, and correlate infrastructure anomalies across network, compute, and service layers at Amazon scale.* Design and develop models for event correlation, root cause analysis, and predictive failure detection using time-series analysis, graph-based methods, and deep learning.* Own the agentic architecture for automated observability workflows, including planning, tool integration, long-horizon reasoning, and multi-agent orchestration.* Define and curate the datasets and evaluation methodologies needed to train, benchmark, and continuously improve detection and classification systems.* Stay deeply hands-on: write production-quality, critical-path code and build core components that take systems from prototype to launch.Partner with Network Development Engineers and operations teams to ground science solutions in real-world infrastructure behavior and operational needs.* Mentor scientists and engineers, raise the science bar, and represent the team in the internal and external scientific community through publications and presentations.A day in the lifeYou will solve real-world problems by analyzing large-scale network telemetry and operational data, designing experiments and simulations, and developing ML models that detect and prevent infrastructure incidents. Your work requires close collaboration with engineers, Network Development Engineers, product managers, and operations leaders across the organization. You will prepare written and verbal presentations to share insights with audiences of varying technical sophistication, and you will iterate rapidly between research and production deployment.Basic qualifications- 3+ years of building machine learning models for business application experience- PhD, or Master's degree and 6+ years of applied research experience- Experience programming in Java, C++, Python or related language- Experience with neural deep learning methods and machine learningPreferred qualification - Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.- Experience in several of the following areas: machine learning, statistics, deep learning, natural language processing, or information retrieval- Experience with time-series analysis, anomaly detection, or graph-based ML methods- Experience with agentic AI architectures, LLMs, or multi-agent systems- Experience applying ML to infrastructure, networking, or observability domainsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 158,800.00 - 214,800.00 USD annually

Vacancy posted 9 days ago
Similar jobs that could be interesting for youBased on the Sr. Applied Scientist, Infrastructure Reliability in Nashville, TN vacancy
  • $114.6k - $234.6k

     ..., and responsible AI. As a Senior Applied Scientist on the team, you will independently own...  ...evaluation protocols and reusable infrastructure. You will examine more than aggregate...  ...contamination, robustness, cost, latency, reliability, safety, and operational constraints.... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Nashville, TN
    4 days ago
  • $169.8k - $355.4k

     ...recognition. Our goal is to enable customers to apply computer vision to solve their...  ...opportunity to work with teams of applied scientists and engineers to deliver high quality...  ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Nashville, TN
    2 days ago
  • $126.2k - $264.1k

     ...production-ready AI solutions. Leads end-to-end applied research across next-generation facial...  ...for real-world deployment. Writes reliable, well-documented code and collaborates...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... 
    Suggested
    Temporary work
    Flexible hours
    Shift work

    Oracle

    Nashville, TN
    14 hours ago
  • $169.3k - $304.7k

     ...generation of the network? Join our Network Infrastructure SRE team! Our team designs, develops...  ...fast, efficient, scalable, and reliable routing software and infrastructure that...  ...financial wellness; Eligibility requirements apply. Equal Employment Opportunity Rights... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Nashville, TN
    2 days ago
  • $142k - $194k

     ...scalable distributed systems or cloud infrastructure. Strong programming skills in Java,...  ...virtualization, storage, networking, security, reliability, or platform automation. You should...  ...and partition-handling policies. We apply load shedding, throttling, and rate... 
    Suggested
    Full time

    Oracle

    Nashville, TN
    2 days ago
  • $146.3k - $306.4k

     ...Establishes KPIs and advanced telemetry; applies formal verification for complex...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...  ...needs within the business unit.System Reliability Performance:Define key performance... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    5 days ago
  •  ...through six business verticals: Testing, Inspection & Consulting; Infrastructure; Utility Services; Environmental Health Sciences; Buildings &...  ...location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not... 
    Senior
    Full time
    Work at office
    Local area
    Long distance
    Night shift
    Weekend work

    NV5, Inc.

    Nashville, TN
    2 days ago
  • $114.6k - $234.6k

     ...elasticity, durability, availability, reliability, and operational-readiness requirements...  ...operational excellence mentoring. Build Infrastructure as Code and operational automation for...  ..., rollbacks, and change management. Apply encryption, access controls, security... 
    Full time
    Work at office
    Relocation package
    Flexible hours

    Oracle

    Nashville, TN
    2 days ago
  • $114.6k - $234.6k

     ...large-scale distributed systems or cloud infrastructure. We need strong experience designing...  ..., performance and load testing, reliability engineering, and production incident response...  ..., and recovery-oriented design. We apply distributed-systems tradeoffs for network... 
    Full time
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle

    Nashville, TN
    12 days ago
  •  ...Infrastructure Architect Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates...  ...issues spanning multiple services, applying advanced investigation and debugging techniques... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Hackajob

    Nashville, TN
    2 days ago
  • $114.6k - $234.6k

    Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises...  ...set the technical direction for their reliability, correctness, security, and operational...  ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for... 
    Temporary work
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle Corporation

    Nashville, TN
    12 days ago
  • $188.1k - $254.5k

    Amazon's Reliability and Maintenance Engineering (RME) Workforce Excellence team seeks an experienced...  ...RME Workforce Excellence team today!The Sr. Manager, Biz Ops will develop and...  ...information. If the country/region you’re applying in isn’t listed, please contact your Recruiting... 
    Senior
    Flexible hours

    Amazon

    Nashville, TN
    3 days ago
  • $114.6k - $234.6k

     ...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting...  ...‑management plans.We are seeking a Core Infrastructure Engineer to design, build, and operate the distributed systems that power reliable, secure, and scalable cloud services.... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $114.6k - $234.6k

     ...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power...  ....System Design & Architecture - System Reliability Design:-Build and design fault-tolerant... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $121.4k - $218.6k

     ...that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...  ...wellness; Eligibility requirements apply. Equal Employment Opportunity Rights... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Nashville, TN
    4 days ago
  • $142.3k - $195.7k

     ...storage architecture and long-term data infrastructure strategy. This role establishes...  ...self-service capabilities that improve reliability and reduce operational effort. Partner...  ...valid job requirements. This policy shall apply to all employment actions, including but... 
    Permanent employment
    Full time
    Temporary work
    Apprenticeship
    Work at office
    Remote work
    Work from home
    Home office

    Humana, Inc.

    Nashville, TN
    4 days ago
  • $142k - $194k

     ...want 3-5+ years of software engineering experience building reliable services, infrastructure, or distributed systems. We need strong programming...  ...using redundancy, replication, and automatic failover. Apply recovery-oriented principles and implement retries,... 
    Senior
    Full time

    Oracle

    Nashville, TN
    25 days ago
  • $123k - $195k

    Job Title Senior Cloud Infrastructure Engineer Job Description You will help maintain and secure...  ...Cloud High services, helping ensure reliable, secure, and compliant cloud operations....  ...Interested candidates are encouraged to apply as soon as possible to ensure consideration... 
    Senior
    Full time
    Work at office
    Immediate start
    Work visa
    Relocation package
    3 days per week

    Philips

    Nashville, TN
    2 days ago
  •  ...involved in executing large-scale, complex infrastructure projects? Connico is a national...  ...through clear communication, sound judgment, reliability, and expert guidance. Participate in...  ...you will also have the opportunity to apply your expertise across an expanding range... 
    Senior
    Full time
    For contractors
    Work at office
    Remote work

    Connico

    Nashville, TN
    a month ago
  • $79.2k - $209.5k

     ...systems, fault tolerance, algorithms, system reliability, and real-time state management. Experience applying machine learning or AI frameworks and techniques...  ...systems that correlate live network and infrastructure signals, detect customer impact, and initiate... 
    Senior
    Full time
    Flexible hours

    Oracle

    Nashville, TN
    11 days ago
  • $79.2k - $209.5k

     ...customer engineering teams and Oracle’s cloud infrastructure, networking, engineering, operations, and support organizations to ensure reliable Day 2 operations, resolve critical issues...  ...expansion of GPU environments.You will apply strong technical judgment to coordinate... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    9 days ago
  • $81.1k - $187k

    Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity...  ...concepts.Patching and Software MaintenanceExperience applying operating system, middleware, and application patches... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    a month ago
  • $79.2k - $209.5k

     ...redundancy, replication, automatic failover), applies recovery‑oriented principles, and...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power...  ....System Design & Architecture - System Reliability Design:-Collaborates with team to build... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $146.3k - $306.4k

     ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...  ....System Design & Architecture - System Reliability Design:Manages the strategy for...  ...expertise in new areas, coaching them to apply learnings to advance the organization.... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    15 days ago
  •  ...energy permitting, with additional opportunities to contribute to onshore wind, electrical transmission, oil & gas, and technology infrastructure projects. As a bat biologist, you will manage bat related tasks including development of proposal scope and budgets, task... 
    Senior
    Remote work
    Night shift

    Environmental Resources Management

    Nashville, TN
    2 days ago
  • $114.6k - $234.6k

     ...building and operating large-scale, highly distributed service infrastructure. You should have experience in an operational environment...  ...simplicity and resilience in mind. We improve service reliability, latency, and operational automation through intelligent tooling... 
    Full time
    Work at office
    Relocation
    Flexible hours

    Oracle

    Nashville, TN
    7 days ago
  • $79.2k - $209.5k

     ...validate, and monitor liquid-cooled GPU infrastructure across factory environments. Develops secure...  ...data-center ingestion, and accelerate reliable hyperscale AI infrastructure delivery....  ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $146.3k - $306.4k

    Oracle Cloud Infrastructure (OCI) is seeking a Lead Principal Core Infrastructure Engineer to join our Networking organization, focusing...  ...architect, build, and enhance solutions to safeguard the reliability and availability of OCI’s global cloud services. The ideal candidate... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    a month ago
  • $19 per hour

     ...an environmental testing laboratory as a Scientist 1, conducting hands-on analysis of air,...  ...supporting the production of accurate, reliable scientific data. This role is suited to...  ...environmental testing. Ability to learn and apply established analytical methods and... 
    Hourly pay
    Full time
    Flexible hours

    Sotalent

    Nashville, TN
    2 days ago
  • $101.55 - $156.73 per hour

     ...Systems (Varian Aria). Conversant in network topology, server infrastructure, backup system design. General programming experience (SQL,...  ...engaged. Learn more about our comprehensive benefits package By applying for a position with Intermountain, I acknowledge that I will... 
    Hourly pay
    Local area
    Remote work
    Monday to Friday

    Intermountain Health

    Nashville, TN
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Applied Scientist, Infrastructure Reliability. Be the first to apply!