Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Applied Scientist, Infrastructure Reliability

$158.8k - $214.8k

Amazon Locker

Are you driven by innovation and complex problem-solving? At Infrastructure Reliability, we build scalable solutions that ensure the reliability of Amazon's critical systems. Our team develops and operates tools that detect and prevent outages to maintain high availability across Amazon's global infrastructure. Join us to architect solutions that directly impact millions of customers, with the resources and support to make meaningful contributions.The team at Amazon is responsible for building intelligent and real-time insights into service-to-service communications, network traffic, and event correlation across hundreds of Amazon's critical fulfillment and robotics services. Our solutions support visibility into anomalous service behavior to prevent and quickly recover from incidents, ensuring high availability to keep the Customer Promise.We are seeking a talented Senior Applied Scientist to invent the next generation of agentic observability solutions at Amazon scale. In this role, you will define, lead, and build the science behind intelligent systems that reason about complex network and infrastructure telemetry, autonomously detect anomalies, and drive automated remediation. You will own the scientific direction end to end, partnering closely with engineering, product, and Network Development Engineers to translate a long-term science vision into concrete research and delivery roadmaps.Working backwards from the needs of our customers and operations teams, you will take the lead on ambiguous, high-impact problems where neither the problem nor the solution is well defined, and deliver production systems that improve infrastructure availability at scale. You will invent new methods, drive their adoption across multiple teams, and remain deeply hands-on with the hardest technical challenges.Key job responsibilities* Define and own the science vision for agentic observability, translating it into research and engineering roadmaps in partnership with product and engineering leaders.* Build ML and agentic AI systems that autonomously detect, classify, and correlate infrastructure anomalies across network, compute, and service layers at Amazon scale.* Design and develop models for event correlation, root cause analysis, and predictive failure detection using time-series analysis, graph-based methods, and deep learning.* Own the agentic architecture for automated observability workflows, including planning, tool integration, long-horizon reasoning, and multi-agent orchestration.* Define and curate the datasets and evaluation methodologies needed to train, benchmark, and continuously improve detection and classification systems.* Stay deeply hands-on: write production-quality, critical-path code and build core components that take systems from prototype to launch.Partner with Network Development Engineers and operations teams to ground science solutions in real-world infrastructure behavior and operational needs.* Mentor scientists and engineers, raise the science bar, and represent the team in the internal and external scientific community through publications and presentations.A day in the lifeYou will solve real-world problems by analyzing large-scale network telemetry and operational data, designing experiments and simulations, and developing ML models that detect and prevent infrastructure incidents. Your work requires close collaboration with engineers, Network Development Engineers, product managers, and operations leaders across the organization. You will prepare written and verbal presentations to share insights with audiences of varying technical sophistication, and you will iterate rapidly between research and production deployment.Basic qualifications- 3+ years of building machine learning models for business application experience- PhD, or Master's degree and 6+ years of applied research experience- Experience programming in Java, C++, Python or related language- Experience with neural deep learning methods and machine learningPreferred qualification - Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.- Experience in several of the following areas: machine learning, statistics, deep learning, natural language processing, or information retrieval- Experience with time-series analysis, anomaly detection, or graph-based ML methods- Experience with agentic AI architectures, LLMs, or multi-agent systems- Experience applying ML to infrastructure, networking, or observability domainsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 158,800.00 - 214,800.00 USD annually

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Sr. Applied Scientist, Infrastructure Reliability in Nashville, TN vacancy
  • $158.3k - $355.4k

    At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for enterprises as a diverse team of creators and inventors. We...  ...company.OCI is seeking an exceptional Senior Principal Applied Scientist to develop AI-enabled platform capabilities for scientific... 
    Senior
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $114.6k - $234.6k

     ..., and responsible AI. As a Senior Applied Scientist on the team, you will independently own...  ...evaluation protocols and reusable infrastructure. You will examine more than aggregate...  ...contamination, robustness, cost, latency, reliability, safety, and operational constraints.... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Nashville, TN
    3 days ago
  • $126.2k - $264.1k

     ...production-ready AI solutions. Leads end-to-end applied research across next-generation facial...  ...for real-world deployment. Writes reliable, well-documented code and collaborates...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... 
    Suggested
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    18 hours ago
  • $169.8k - $355.4k

     ...recognition. Our goal is to enable customers to apply computer vision to solve their...  ...opportunity to work with teams of applied scientists and engineers to deliver high quality...  ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... 
    Senior
    Temporary work
    Flexible hours

    Oracle

    Nashville, TN
    1 day ago
  • $141.2k - $191.1k

    The Technical Infrastructure Program Manager III (TIPM III) within Data Center Infrastructure Engineering...  ..., and field teams to deliver scalable, reliable, and cost-effective infrastructure. This...  ...employment. The benefits that generally apply to regular, full-time employees include:... 
    Senior
    Full time
    Temporary work
    Seasonal work
    Worldwide
    Flexible hours
    Night shift
    Day shift

    Amazon

    Nashville, TN
    2 days ago
  • $79.2k - $209.5k

    As a Senior Core Infrastructure Engineer, you will design and develop major new features within...  ....System Design & Architecture - System Reliability Design:-Collaborates with team to build...  ..., and automatic failover mechanisms.-Applies recovery oriented computing principles... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    18 hours ago
  • $146.3k - $306.4k

     ...distributed-systems, developer-platform, and reliability problems with broad impact across OCI....  ...who are excited by large-scale cloud infrastructure, deployment orchestration, operational...  ..., and alerting for complex systems.Apply formal verification techniques such as... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $146.3k - $306.4k

     ...Establishes KPIs and advanced telemetry; applies formal verification for complex...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...  ...needs within the business unit.System Reliability Performance:Define key performance... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $114.6k - $234.6k

     ...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting...  ...‑management plans.We are seeking a Core Infrastructure Engineer to design, build, and operate the distributed systems that power reliable, secure, and scalable cloud services.... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $142k - $194k

     ...scalable distributed systems or cloud infrastructure. Strong programming skills in Java,...  ...virtualization, storage, networking, security, reliability, or platform automation. You should...  ...and partition-handling policies. We apply load shedding, throttling, and rate... 
    Full time

    Oracle

    Nashville, TN
    6 days ago
  •  ...Infrastructure Architect Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates...  ...issues spanning multiple services, applying advanced investigation and debugging techniques... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Hackajob

    Nashville, TN
    1 day ago
  •  ...through six business verticals: Testing, Inspection & Consulting; Infrastructure; Utility Services; Environmental Health Sciences; Buildings &...  ...location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not... 
    Senior
    Full time
    Work at office
    Local area
    Long distance
    Night shift
    Weekend work

    NV5, Inc.

    Nashville, TN
    1 day ago
  • $114.6k - $234.6k

     ...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power...  ....System Design & Architecture - System Reliability Design:-Build and design fault-tolerant... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $114.6k - $234.6k

    Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises...  ...set the technical direction for their reliability, correctness, security, and operational...  ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for... 
    Temporary work
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $142k - $194k

     ...close security gaps. We ensure cloud infrastructure complies with relevant industry standards...  ...and industry practices. We seek and apply feedback to improve performance and coach...  ...the distributed systems that power reliable, secure, and scalable cloud services. This... 
    Full time

    Oracle

    Nashville, TN
    18 days ago
  • $79.2k - $209.5k

     ...customer engineering teams and Oracle’s cloud infrastructure, networking, engineering, operations, and support organizations to ensure reliable Day 2 operations, resolve critical issues...  ...expansion of GPU environments.You will apply strong technical judgment to coordinate... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $123k - $195k

     ...Job TitleSenior Cloud Infrastructure EngineerJob DescriptionYou will help maintain and secure...  ...Community Cloud High services, helping ensure reliable, secure, and compliant cloud operations....  ...Interested candidates are encouraged to apply as soon as possible to ensure... 
    Senior
    Full time
    Work at office
    Immediate start
    Work visa
    Relocation package
    3 days per week

    Philips

    Nashville, TN
    4 days ago
  • $146.3k - $306.4k

     ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...  ....System Design & Architecture - System Reliability Design:Manages the strategy for...  ...expertise in new areas, coaching them to apply learnings to advance the organization.... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $146.3k - $306.4k

     ...technical and business impact.The Oracle Cloud Infrastructure (OCI) team can provide you the...  ...software engineering division, you will apply your knowledge of software architecture...  ...that is highly performant, scalable, and reliable.Ownership Scope - You own the performance... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    18 hours ago
  • $114.6k - $234.6k

     ....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...  ...on data center and WAN orchestration.Apply deep expertise in networking...  ...distributed systems concepts, and system reliability.Demonstrated experience building resilient... 
    Temporary work
    Worldwide
    Flexible hours

    Oracle Corporation

    Nashville, TN
    9 days ago
  • $114.6k - $234.6k

     ...large-scale distributed systems or cloud infrastructure. We need strong experience designing...  ..., performance and load testing, reliability engineering, and production incident response...  ..., and recovery-oriented design. We apply distributed-systems tradeoffs for network... 
    Full time
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle

    Nashville, TN
    6 days ago
  • $81.1k - $187k

    Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity...  ...concepts.Patching and Software MaintenanceExperience applying operating system, middleware, and application patches... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $79.2k - $209.5k

     ...Staff, you will help build Lightweight Infrastructure (LWI), the next-generation runtime and...  ...software engineering experience building reliable services, infrastructure, or...  ...redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements... 
    Senior
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $81.1k - $187k

    Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity...  ...files to identify and resolve problems. Apply operating system, middleware, and application patches... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $142k - $194k

     ...want 3-5+ years of software engineering experience building reliable services, infrastructure, or distributed systems. We need strong programming...  ...using redundancy, replication, and automatic failover. Apply recovery-oriented principles and implement retries,... 
    Senior
    Full time

    Oracle

    Nashville, TN
    19 days ago
  • $79.2k - $209.5k

     ...or related experience.The Oracle Cloud Infrastructure (OCI) team can provide you the opportunity...  ...biggest challenges for the team are reliability and performance. The growth of the business...  ...features and work plans.Expertise applying threat modeling and other risk-identification... 
    Senior
    Temporary work
    Live in
    Remote work
    Relocation package
    Long distance
    Flexible hours
    Night shift

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $1,500 - $1,730 per week

     ...The most jobs in the industry. We have the largest and most reliable job database, which means the jobs you see are open, updated in...  ...wages and tax-free expense reimbursements. Aya is an Equal Employment Opportunity ("EEO") Employer and welcomes all to apply.... 
    Weekly pay
    Daily paid
    Permanent employment
    Full time
    Contract work
    Local area
    Immediate start
    Relocation
    Shift work
    Night shift

    Aya Healthcare

    Nashville, TN
    21 hours ago
  • $6,040 - $6,201 per week

     ...The most jobs in the industry. We have the largest and most reliable job database, which means the jobs you see are open, updated in...  ...wages and tax-free expense reimbursements. Aya is an Equal Employment Opportunity ("EEO") Employer and welcomes all to apply.... 
    Weekly pay
    Daily paid
    Permanent employment
    Full time
    Contract work
    Local area
    Immediate start
    Relocation
    Shift work

    Aya Healthcare

    Nashville, TN
    1 day ago
  • $1,233 - $1,417 per week

     ...ensuring accurate labeling and handling Ensure accuracy and reliability of test results by following quality control protocols Follow...  ...with Fusion Medical Staffing and join our mission to improve lives. Apply now! *Fusion is an EOE/E-Verify Employer... 
    Full time
    Contract work
    Temporary work
    Immediate start
    Shift work
    Night shift

    Fusion Medical Staffing

    Nashville, TN
    21 hours ago
  •  ...involved in executing large-scale, complex infrastructure projects? Connico is a national...  ...through clear communication, sound judgment, reliability, and expert guidance. Participate in...  ...you will also have the opportunity to apply your expertise across an expanding range... 
    Senior
    Full time
    For contractors
    Work at office
    Remote work

    Connico

    Nashville, TN
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Applied Scientist, Infrastructure Reliability. Be the first to apply!