Sr. Applied Scientist, Infrastructure Reliability
$158.8k - $214.8kAmazon Locker
Are you driven by innovation and complex problem-solving? At Infrastructure Reliability, we build scalable solutions that ensure the reliability of Amazon's critical systems. Our team develops and operates tools that detect and prevent outages to maintain high availability across Amazon's global infrastructure. Join us to architect solutions that directly impact millions of customers, with the resources and support to make meaningful contributions.The team at Amazon is responsible for building intelligent and real-time insights into service-to-service communications, network traffic, and event correlation across hundreds of Amazon's critical fulfillment and robotics services. Our solutions support visibility into anomalous service behavior to prevent and quickly recover from incidents, ensuring high availability to keep the Customer Promise.We are seeking a talented Senior Applied Scientist to invent the next generation of agentic observability solutions at Amazon scale. In this role, you will define, lead, and build the science behind intelligent systems that reason about complex network and infrastructure telemetry, autonomously detect anomalies, and drive automated remediation. You will own the scientific direction end to end, partnering closely with engineering, product, and Network Development Engineers to translate a long-term science vision into concrete research and delivery roadmaps.Working backwards from the needs of our customers and operations teams, you will take the lead on ambiguous, high-impact problems where neither the problem nor the solution is well defined, and deliver production systems that improve infrastructure availability at scale. You will invent new methods, drive their adoption across multiple teams, and remain deeply hands-on with the hardest technical challenges.Key job responsibilities* Define and own the science vision for agentic observability, translating it into research and engineering roadmaps in partnership with product and engineering leaders.* Build ML and agentic AI systems that autonomously detect, classify, and correlate infrastructure anomalies across network, compute, and service layers at Amazon scale.* Design and develop models for event correlation, root cause analysis, and predictive failure detection using time-series analysis, graph-based methods, and deep learning.* Own the agentic architecture for automated observability workflows, including planning, tool integration, long-horizon reasoning, and multi-agent orchestration.* Define and curate the datasets and evaluation methodologies needed to train, benchmark, and continuously improve detection and classification systems.* Stay deeply hands-on: write production-quality, critical-path code and build core components that take systems from prototype to launch.Partner with Network Development Engineers and operations teams to ground science solutions in real-world infrastructure behavior and operational needs.* Mentor scientists and engineers, raise the science bar, and represent the team in the internal and external scientific community through publications and presentations.A day in the lifeYou will solve real-world problems by analyzing large-scale network telemetry and operational data, designing experiments and simulations, and developing ML models that detect and prevent infrastructure incidents. Your work requires close collaboration with engineers, Network Development Engineers, product managers, and operations leaders across the organization. You will prepare written and verbal presentations to share insights with audiences of varying technical sophistication, and you will iterate rapidly between research and production deployment.Basic qualifications- 3+ years of building machine learning models for business application experience- PhD, or Master's degree and 6+ years of applied research experience- Experience programming in Java, C++, Python or related language- Experience with neural deep learning methods and machine learningPreferred qualification - Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.- Experience in several of the following areas: machine learning, statistics, deep learning, natural language processing, or information retrieval- Experience with time-series analysis, anomaly detection, or graph-based ML methods- Experience with agentic AI architectures, LLMs, or multi-agent systems- Experience applying ML to infrastructure, networking, or observability domainsAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, TN, Nashville - 158,800.00 - 214,800.00 USD annually
$158.3k - $355.4k
At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for enterprises as a diverse team of creators and inventors. We... ...company.OCI is seeking an exceptional Senior Principal Applied Scientist to develop AI-enabled platform capabilities for scientific...SeniorTemporary workImmediate startFlexible hours$114.6k - $234.6k
..., and responsible AI. As a Senior Applied Scientist on the team, you will independently own... ...evaluation protocols and reusable infrastructure. You will examine more than aggregate... ...contamination, robustness, cost, latency, reliability, safety, and operational constraints....SeniorTemporary workFlexible hours$126.2k - $264.1k
...production-ready AI solutions. Leads end-to-end applied research across next-generation facial... ...for real-world deployment. Writes reliable, well-documented code and collaborates... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power...SuggestedTemporary workFlexible hoursShift work$169.8k - $355.4k
...recognition. Our goal is to enable customers to apply computer vision to solve their... ...opportunity to work with teams of applied scientists and engineers to deliver high quality... ...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything...SeniorTemporary workFlexible hours$141.2k - $191.1k
The Technical Infrastructure Program Manager III (TIPM III) within Data Center Infrastructure Engineering... ..., and field teams to deliver scalable, reliable, and cost-effective infrastructure. This... ...employment. The benefits that generally apply to regular, full-time employees include:...SeniorFull timeTemporary workSeasonal workWorldwideFlexible hoursNight shiftDay shift$79.2k - $209.5k
As a Senior Core Infrastructure Engineer, you will design and develop major new features within... ....System Design & Architecture - System Reliability Design:-Collaborates with team to build... ..., and automatic failover mechanisms.-Applies recovery oriented computing principles...SeniorTemporary workFlexible hours$146.3k - $306.4k
...distributed-systems, developer-platform, and reliability problems with broad impact across OCI.... ...who are excited by large-scale cloud infrastructure, deployment orchestration, operational... ..., and alerting for complex systems.Apply formal verification techniques such as...Temporary workFlexible hours$146.3k - $306.4k
...Establishes KPIs and advanced telemetry; applies formal verification for complex... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... ...needs within the business unit.System Reliability Performance:Define key performance...Temporary workFlexible hours$114.6k - $234.6k
...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting... ...‑management plans.We are seeking a Core Infrastructure Engineer to design, build, and operate the distributed systems that power reliable, secure, and scalable cloud services....Temporary workFlexible hours$142k - $194k
...scalable distributed systems or cloud infrastructure. Strong programming skills in Java,... ...virtualization, storage, networking, security, reliability, or platform automation. You should... ...and partition-handling policies. We apply load shedding, throttling, and rate...Full time- ...Infrastructure Architect Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates... ...issues spanning multiple services, applying advanced investigation and debugging techniques...Temporary workImmediate startFlexible hoursShift work
- ...through six business verticals: Testing, Inspection & Consulting; Infrastructure; Utility Services; Environmental Health Sciences; Buildings &... ...location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not...SeniorFull timeWork at officeLocal areaLong distanceNight shiftWeekend work
$114.6k - $234.6k
...failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting... ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... ....System Design & Architecture - System Reliability Design:-Build and design fault-tolerant...Temporary workFlexible hoursShift work$114.6k - $234.6k
Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises... ...set the technical direction for their reliability, correctness, security, and operational... ...failover, and recovery-oriented design.Apply sound distributed-systems tradeoffs for...Temporary workWork at officeWorldwideRelocationRelocation packageFlexible hours$142k - $194k
...close security gaps. We ensure cloud infrastructure complies with relevant industry standards... ...and industry practices. We seek and apply feedback to improve performance and coach... ...the distributed systems that power reliable, secure, and scalable cloud services. This...Full time$79.2k - $209.5k
...customer engineering teams and Oracle’s cloud infrastructure, networking, engineering, operations, and support organizations to ensure reliable Day 2 operations, resolve critical issues... ...expansion of GPU environments.You will apply strong technical judgment to coordinate...SeniorTemporary workFlexible hours$123k - $195k
...Job TitleSenior Cloud Infrastructure EngineerJob DescriptionYou will help maintain and secure... ...Community Cloud High services, helping ensure reliable, secure, and compliant cloud operations.... ...Interested candidates are encouraged to apply as soon as possible to ensure...SeniorFull timeWork at officeImmediate startWork visaRelocation package3 days per week$146.3k - $306.4k
...Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... ....System Design & Architecture - System Reliability Design:Manages the strategy for... ...expertise in new areas, coaching them to apply learnings to advance the organization....SeniorTemporary workFlexible hours$146.3k - $306.4k
...technical and business impact.The Oracle Cloud Infrastructure (OCI) team can provide you the... ...software engineering division, you will apply your knowledge of software architecture... ...that is highly performant, scalable, and reliable.Ownership Scope - You own the performance...SeniorTemporary workFlexible hours$114.6k - $234.6k
....Only Oracle brings together the data, infrastructure, applications, and expertise to power everything... ...on data center and WAN orchestration.Apply deep expertise in networking... ...distributed systems concepts, and system reliability.Demonstrated experience building resilient...Temporary workWorldwideFlexible hours$114.6k - $234.6k
...large-scale distributed systems or cloud infrastructure. We need strong experience designing... ..., performance and load testing, reliability engineering, and production incident response... ..., and recovery-oriented design. We apply distributed-systems tradeoffs for network...Full timeWork at officeWorldwideRelocationRelocation packageFlexible hours$81.1k - $187k
Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity... ...concepts.Patching and Software MaintenanceExperience applying operating system, middleware, and application patches...SeniorTemporary workFlexible hours$79.2k - $209.5k
...Staff, you will help build Lightweight Infrastructure (LWI), the next-generation runtime and... ...software engineering experience building reliable services, infrastructure, or... ...redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements...SeniorTemporary workImmediate startFlexible hours$81.1k - $187k
Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity... ...files to identify and resolve problems. Apply operating system, middleware, and application patches...SeniorTemporary workFlexible hours$142k - $194k
...want 3-5+ years of software engineering experience building reliable services, infrastructure, or distributed systems. We need strong programming... ...using redundancy, replication, and automatic failover. Apply recovery-oriented principles and implement retries,...SeniorFull time$79.2k - $209.5k
...or related experience.The Oracle Cloud Infrastructure (OCI) team can provide you the opportunity... ...biggest challenges for the team are reliability and performance. The growth of the business... ...features and work plans.Expertise applying threat modeling and other risk-identification...SeniorTemporary workLive inRemote workRelocation packageLong distanceFlexible hoursNight shift$1,500 - $1,730 per week
...The most jobs in the industry. We have the largest and most reliable job database, which means the jobs you see are open, updated in... ...wages and tax-free expense reimbursements. Aya is an Equal Employment Opportunity ("EEO") Employer and welcomes all to apply....Weekly payDaily paidPermanent employmentFull timeContract workLocal areaImmediate startRelocationShift workNight shift$6,040 - $6,201 per week
...The most jobs in the industry. We have the largest and most reliable job database, which means the jobs you see are open, updated in... ...wages and tax-free expense reimbursements. Aya is an Equal Employment Opportunity ("EEO") Employer and welcomes all to apply....Weekly payDaily paidPermanent employmentFull timeContract workLocal areaImmediate startRelocationShift work$1,233 - $1,417 per week
...ensuring accurate labeling and handling Ensure accuracy and reliability of test results by following quality control protocols Follow... ...with Fusion Medical Staffing and join our mission to improve lives. Apply now! *Fusion is an EOE/E-Verify Employer...Full timeContract workTemporary workImmediate startShift workNight shift- ...involved in executing large-scale, complex infrastructure projects? Connico is a national... ...through clear communication, sound judgment, reliability, and expert guidance. Participate in... ...you will also have the opportunity to apply your expertise across an expanding range...SeniorFull timeFor contractorsWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Applied Scientist, Infrastructure Reliability. Be the first to apply!
- regulatory scientist Nashville, TN
- graduate scientist Nashville, TN
- manufacturing scientist Nashville, TN
- analytical scientist Nashville, TN
- support scientist Nashville, TN
- application scientist Nashville, TN
- drug safety scientist Nashville, TN
- machine learning scientist Nashville, TN
- lab scientist Nashville, TN
- safety scientist Nashville, TN



