Site Reliability Engineer (SRE) - AI Platform & Cloud
Morgan Stanley
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.
This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.
Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
What you'll do in the role:
Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
Design and build automation for core platform capabilities, reducing manual toil
Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
Optimize cost vs. performance tradeoffs in large-scale compute environments
Harden systems for security, compliance, auditability, and data governance
Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
Maintain runbooks, operational playbooks, documentation, and training materials
Participate in on-call rotations and respond to production incidents 24/7 as needed
Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
What you ' ll bring to the role:
Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
5 years of production experience in SRE / Infrastructure / ops for large-scale systems
Strong programming/scripting skills (Python, Go, Java, or equivalent)
Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
Solid experience in capacity planning, performance tuning, scaling, and incident response
Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
Excellent communication, documentation, and cross-team collaboration skills
Proven track record of reducing operational toil via automation
Nice to have
Understanding of SRE techniques.
Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
Understanding of ModelOps/ ML Ops/ LLM Op.
Experience with chaos engineering, canary deployments, blue/green rollouts
We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.
All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren’t just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.
Wherever you are in our 1,200 global offices, you’ll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.
At Morgan Stanley Alpharetta, we support the Firm’s global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.
Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.
It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.
Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit: .
- ...leading workflow management platform built exclusively for K-12... ...the future of education. Site Reliability Engineer (SRE) Overview: We are... ...just operate a dashboard. AI-Accelerated Execution (core... ...Automation, Infrastructure & Cloud: Proficient in Python, Go,...PlatformCloudFull timeLive inWork at office
$125k - $175k
...communities. This is an SRE/Production Support... ...the operational reliability of deployed... ...Management application platforms, participation in... ...Science, Computer Engineering). - 5+ years’... ...in AWS/GCP/Azure Cloud technologies - Experience... ...- Hands-on with AI and implementation...PlatformCloudFull timeTemporary work- ...Job Description Job Description We are hiring an SRE Platform Engineer (IBM BPM/ODM) with our partner in Alpharetta, GA for an onsite role... ...the division such as operating systems, public and private cloud, web, and middleware components like Business Process Management...PlatformCloudLocal area
- ...forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of... ...: 6+ years as an SRE, DevOps Engineer, or similar role Cloud: Strong experience with AWS (EKS, EC2, S3, Route53, IAM)...CloudPermanent employmentWork experience placementLocal area
$129k - $161k
...Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20... ...engineering experience, including 3+ years in SRE or reliability-focused roles. ~ Demonstrated...PlatformCloudRemote work- ...Java Forward AI Developer Alpharetta, GA (Onsite) – ONLY LOCAL TO ALPHARETTA, GA Required... ...2EE / Spring Boot. Strong understanding of SRE principles and practices. Hands-on experience with Google Cloud Platform (GCP). Experience with cloud-native...PlatformCloudTemporary workLocal area
- ...empowered to shape the future of payments. As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE,... ..., observability, and operational excellence across cloud platforms. Key Responsibilities Define and manage SLOs/SLIs...PlatformCloudFull timeWorldwide
$129.5k - $186.1k
Staff Platform Engineer At UKG, the work you do matters. The code... ...on-premises, private cloud, and public cloud environments... ...who enjoys building reliable platforms, automating... ...engineering, DevOps, SRE, or infrastructure engineering... ..., and people-first AI, our ability to reveal...PlatformCloud- ...The Site Reliability Engineer (Observability) to join its Platform Engineering organization and help develop, implement, and mature... ..., Kubernetes environments, cloud platforms, and Windows/Linux systems... ...ideal candidate combines strong SRE/DevOps engineering experience...PlatformCloud
- ...scripting is a plus Provisioning cloud infrastructure (AWS, GCP) using... ...Collaborating with product, architecture, and engineering groups to build a platform that streamlines application... ...level experience in DevOps, Site Reliability Engineering with expertise in Enterprise...PlatformCloud
$95k - $130k
...) Operations team as a VM Ops Analyst in Cloud Platform Security and Developer Enablement to assist... ...security operations, or related SRE functions.Vulnerability management / Cybersecurity... ...:Basic understanding of emerging LLM/AI cyber threatsAbility to document user-stories...PlatformCloudTemporary workWeekend work$61.07 per hour
...Our client, a technology and cloud operations organization is seeking a Senior Site Reliability Engineer to join their team. As a Senior Site Reliability Engineer, you... ...Needed? ~4+ years of experience with the Azure platform. ~4+ years of practical experience...PlatformCloudWeekly payTemporary workWork experience placementFlexible hours- Machine Learning Engineer / Data Scientist We are seeking a skilled and... ...with expertise in Google Cloud Platform (GCP), Databricks, Kubernetes, and Generative AI. This role involves designing... ...systems to ensure performance and reliability in production environments. Key...PlatformCloud
$92.16k - $138.24k
...currently seeking a Backend & AI/ML Engineer to join our team in Alpharetta... ...Job Duties: Develop and support cloud-native data engineering solutions on Google Cloud Platform.Build and maintain data... ...to NTT DATA offices or client sites. This ensures we can provide timely...PlatformCloudFull timeTemporary workWork at officeRemote workFlexible hours- ...OEConnection LLC is seeking a Senior AI Platform Engineer to design, build, and evolve a cloud platform powering engineering teams. You will develop scalable... ...with IaC, and optimize multi-cloud environments for reliability and cost efficiency. You will collaborate with...PlatformCloud
- ...Learning, and Generative AI technologies. The... ...frameworks, APIs, and cloud‑native architectures.... ...Gemini, Claude, or similar platforms. Develop data... ...Owners, Architects, Data Engineers, and Business Stakeholders... ..., security, and reliability. Participate in Agile...PlatformCloud
- ...that support scalable, reliable, and maintainable solutions... ...Salesforce, ERP platforms, data warehouses, marketing... ....Apply software engineering best practices, including... ...Salesforce CPQ, Experience Cloud, Marketing Cloud, Data... ...artificial intelligence (AI) tools to support parts...PlatformCloudFull time
- ...Salesforce Developer - Agentforce (Salesforce AI-driven solutions and automation) -... ...with strong expertise in the Salesforce platform and hands-on experience with Agentforce (... ...implement solutions on Salesforce (Sales Cloud, Service Cloud, Experience Cloud) · Build...PlatformCloudTemporary work
$155.4k - $261.1k
Principal Software Engineer - (Kubernetes - AKS)Join AT... ...enterprise forward. Our Systems Reliability and Software Delivery... ..., we are expanding our platform engineering... ...design, build, and enhance cloud-native platform capabilities... ...etc), PackerStrong AI assisted delivery mindsetHands...PlatformCloudTemporary workWork at officeLocal areaRelocation- ...high availability and reliability.Mentor junior developers... ...in Computer Science, Engineering, or equivalent work... ...CI/CD).Experience with cloud platforms (AWS, Azure, or GCP) is... ...of the world's leading AI and digital... ...DATA offices or client sites. This ensures we can provide...PlatformCloudFull timeTemporary workWork at officeRemote workFlexible hours
- ...~4+ years of experience in an SRE, DevOps, or cloud infrastructure role. ~ Strong experience with Azure cloud services and infrastructure. ~ Hands-on experience with java and Terraform and Terragrunt for infrastructure-as-code. ~ Proficiency with Kubernetes (preferably...Cloud
- ...Doeren Mayhew is seeking a Senior Software Engineer . This position is available in Dallas,... ...Identify opportunities to introduce AI driven capabilities, automation, and data... ...and SQL optimization ~ Exposure to cloud platforms (Azure, AWS, or GCP) and modern deployment...PlatformCloud
- ...Equifax seeks a visionary full stack engineer to join our technology transformation. You will design and deploy cloud-native APIs, microservices, and PaaS/SaaS platforms across GCP/AWS, applying AI-assisted coding tools to accelerate delivery. You will lead a cross...PlatformCloud
$196.1k - $326.9k
...Responsibilities: Proven leadership of AI/ML engineering orgs through a management... ...Product, Data Engineering, Platform/Infra, Security, Legal/... ...leadersExperience with enterprise-scale cloud AI/ML platforms (Azure... ...are posted on our career site: careers.mckesson.com.McKesson...PlatformCloudFull time- ...seeking an IT Support Engineer who thrives in a communicative... ...to ensure smooth and reliable service delivery. This... ..., permissions, and site structure Microsoft... ...Exposure to cloud platforms such as Azure or AWS... ...Microsoft Copilot or similar AI-assist tools Certifications...PlatformCloud
- ...We are seeking a Software Engineer to join our backend engineering... ...next generation of LoadUp's cloud-native platform. This role focuses on... ...development tools, including AI-assisted workflows, to increase... ...maintaining high quality and reliability. Responsibilities Backend...PlatformCloudHourly payContract workLocal area
$162.4k - $211.9k
...JOB TITLE: Lead System Engineering JOB LOCATION: 500 North Point... ...solution (consisting of platform, network, software, cloud, etc.) through functional, performance, and reliability analysis using engineering... ...Enterprise Architect. Utilize AI and Machine Learning (ML) technologies...PlatformCloudTemporary workLocal area- Software Engineer The Software Engineer builds platforms—delivering high‐quality.NET solutions, secure APIs, and cloud‐ready features that improve efficiency for internal teams, clients, and... ...applications, integrate modern AI capabilities, and mentor junior engineers...PlatformCloud
- ...Francisco Partners is seeking a Senior AI Platform Engineer to design, build, and evolve a scalable cloud platform for AI-powered development. You will work with... ...to improve developer productivity and platform reliability. Responsibilities include IaC, cloud management...PlatformCloud
- ...organizations — helping them harness AI to drive outcomes at a time of... ...Analytics CloudPayment platforms and banking systemsPartner with... ...Connectivity.Experience with cloud-based treasury and banking solutions... ...ability to access or use this site as a result of your disability...PlatformCloudMinimum wageFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Platform & Cloud. Be the first to apply!
- site reliability engineer Alpharetta, GA
- platform engineer Alpharetta, GA
- platform developer Alpharetta, GA
- senior cloud data engineer Alpharetta, GA
- cloud engineer Alpharetta, GA
- big data cloud engineer Alpharetta, GA
- aws cloud architect Alpharetta, GA
- aws cloud security engineer Alpharetta, GA
- cloud developer Alpharetta, GA
- informatica cloud developer Alpharetta, GA




