Site Reliability Engineer (SRE) - AI Platform & Cloud
Morgan Stanley
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.
This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.
Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
What you'll do in the role:
Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
Design and build automation for core platform capabilities, reducing manual toil
Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
Optimize cost vs. performance tradeoffs in large-scale compute environments
Harden systems for security, compliance, auditability, and data governance
Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
Maintain runbooks, operational playbooks, documentation, and training materials
Participate in on-call rotations and respond to production incidents 24/7 as needed
Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
What you ' ll bring to the role:
Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
5 years of production experience in SRE / Infrastructure / ops for large-scale systems
Strong programming/scripting skills (Python, Go, Java, or equivalent)
Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
Solid experience in capacity planning, performance tuning, scaling, and incident response
Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
Excellent communication, documentation, and cross-team collaboration skills
Proven track record of reducing operational toil via automation
Nice to have
Understanding of SRE techniques.
Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
Understanding of ModelOps/ ML Ops/ LLM Op.
Experience with chaos engineering, canary deployments, blue/green rollouts
We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.
All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren’t just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.
Wherever you are in our 1,200 global offices, you’ll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.
At Morgan Stanley Alpharetta, we support the Firm’s global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.
Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.
It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.
Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit: .
- ...leading workflow management platform built exclusively for K-12... ...the future of education. Site Reliability Engineer (SRE) Overview: We are... ...just operate a dashboard. AI-Accelerated Execution (core... ...Automation, Infrastructure & Cloud: Proficient in Python, Go,...PlatformCloudFull timeLive inWork at office
$125k - $175k
...communities. This is an SRE/Production Support... ...the operational reliability of deployed... ...Management application platforms, participation in... ...Science, Computer Engineering). - 5+ years’... ...in AWS/GCP/Azure Cloud technologies - Experience... ...- Hands-on with AI and implementation...PlatformCloudFull timeTemporary work$86.6k - $144.4k
...building secure, resilient cloud platforms and using automation to... ...solve complex security and reliability challenges?Do you enjoy... ...compliance-as-code, and AI-driven innovation?About the... ...Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role...PlatformCloudFull timeLocal area- ...Production Management & Reliability Engineering position at Director... ...organization delivers platforms that support core client... ...We are looking for a Site Reliability Engineer... ...Python, Perl, etc.) and cloud driven development ~... ...leveraging generative AI tools to enhance...PlatformCloudFull timeWork visaFlexible hoursWeekend work
$125k - $175k
...communities. This is a Lead Site Reliability Engineer position at Vice... ...Management application platforms, participation in key... ...Engineering PracticesChampion SRE principles, including... ...capabilities, and AI-driven operational... ...Python, Perl, etc.) and cloud driven...PlatformCloudTemporary work- ...is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting... ...customer- and partner-facing platforms that power payment... ...engineering standards for multi-cloud infrastructure across AWS... ...artificial intelligence (AI) tools to support parts...PlatformCloud
$129k - $161k
...Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20... ...engineering experience, including 3+ years in SRE or reliability-focused roles. ~ Demonstrated...PlatformCloudRemote work- ...On-Site role Job Description: ~4 - 1... ...pm) ~ Database Site Reliability Engineer (Database Operations)... ...Reliability Engineer (SRE) to support and operate... ...mission-critical database platforms within a fast-paced... ..., virtual, and cloud environments. Implement...PlatformCloudPermanent employmentTemporary work
$86.6k - $144.4k
...professional software engineering experience (US... ...developing and supporting cloud-native... ...available enterprise platforms and participating... ...architectures. Quality, Reliability, and Operations... ...leveraging AI-assisted development... ...Product, Architecture, SRE, QA, and peer...PlatformCloudLocal area- ...Java Forward AI Developer Alpharetta, GA (Onsite) – ONLY LOCAL TO ALPHARETTA, GA Required... ...2EE / Spring Boot. Strong understanding of SRE principles and practices. Hands-on experience with Google Cloud Platform (GCP). Experience with cloud-native...PlatformCloudTemporary workLocal area
- ...Learning, and Generative AI technologies. The... ...frameworks, APIs, and cloud-native architectures.... ...Gemini, Claude, or similar platforms. • Develop data... ...Owners, Architects, Data Engineers, and Business Stakeholders... ..., security, and reliability. • Participate in Agile...PlatformCloudFull time
- ...IT Support Engineer We are a growing organization... ...to ensure smooth and reliable service delivery. This... ...governance, permissions, and site structure Microsoft... ...Exposure to cloud platforms such as Azure or AWS... ...Microsoft Copilot or similar AI-assist tools Certifications...PlatformCloud
- ...Exposure to data preprocessing tools (Pandas, NumPy) - Basic cloud AI services Mid Level (5–12 years) Mandatory Skills: -... ...Skills: - Experience with NLP, deep learning - Cloud ML platforms (AWS SageMaker, Azure AI) Expert (12+ years) Mandatory...PlatformCloud
- ...Sr. Software Engineer Mlops Engineer Location: NYC/Alpharetta... ...Strong understanding of AI/Client concepts and techniques... ...Familiarity with at least one cloud platforms / hyper-scalers (e.g., GCP,... ...scalability, performance, and reliability. Collaborate to integrate...PlatformCloud
- ...ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of... ...certificationFamiliarity with .NET application stackMulti-cloud exposureExperience managing Kubernetes clusters with Rancher...CloudPermanent employmentFull timeWork experience placementLocal area
- ...Learning, Java, data warehousing, cloud computing, and build engineering. As a global provider of IT... ...in Python, Node.js, cloud platforms (AWS/Azure), and experience integrating AI/ML capabilities into... ...~ Maintain high system reliability, uptime, and performance standards...PlatformCloudShift work
- ...Job Title: Python Engineer Location: Alpharetta, GA... ...• Solid understanding of AI/ML workflows: training, validation... ...Experience deploying models to cloud platforms (AWS, GCP, or Azure). •... ...and contract tests to ensure reliability and stability. • Create and...PlatformCloudFull timeContract work
- ...closely with business stakeholders, data engineers, and product teams to develop predictive... ...BI, or Matplotlib. Knowledge of cloud platforms such as AWS, Azure, or GCP. Experience... ...domains. Knowledge of Generative AI, LLMs, NLP, and MLOps frameworks. Experience...PlatformCloud
$92.6k - $173.1k
...deployment of advanced generative AI models, including large... ...Analyzing the latest trends such as cloud computing and distributed... ...LLM fine-tuning and prompt engineering. ~ Proficiency with data wrangling... ...computing, GPU, cloud platforms, and Big Data domains (e.g.,...PlatformCloudSummer holidayLocal areaFlexible hours- ...Job Title: AI/ML Engineer Location : Alpharetta GA Interview type: final discussion-... ...preprocessing tools (Pandas, NumPy) Basic cloud AI services Mid Level (5-12 years)... ...with NLP, deep learning Cloud ML platforms (AWS SageMaker, Azure AI) Expert (12...PlatformCloudHourly payContract work
$128k - $216k
...times a day - quickly, reliably, and securely. Any... ....Job TitleSr. Site Reliability EngineerAbout... ...management platform that handles everything... ...Senior Site Reliability Engineer do at Fiserv?As a... ...greatest of what cloud technology at... ...Reviews, mentor teams on SRE principles, and...PlatformCloudFull timeWorldwide- ...Machine Learning Engineer Equifax is where you can power your... ...opportunities and limitations of AI, ML, and data engineering... ...practices, scalability and reliability ~ Cloud Certification Strongly Preferred... ...through our online learning platform with guided career tracks....PlatformCloudWork experience placementImmediate start
- ...you a progressive software engineer, an advocate of agile development... ...team is at the forefront of AI-assisted software development... ...the variety of environments, platforms, technologies & languages, you... ..., web services and hybrid cloud environments. Primary Responsibilities...PlatformCloud
- ...classification, data extraction) by orchestrating OCR/AI services from Pega workflows and managing... ...DB Trace, and PDC (Predictive Diagnostic Cloud); tune data pages, query strategies,... .... Conduct impact assessments for platform upgrades (Pega version/hotfixes), ruleset...PlatformCloud
- ...Senior Software Engineer Driven by transformative digital technologies... ...RIB North America's SpecLink platform team, a collaborative and forward-thinking group building a cloud-native SaaS product for the... ...practices ~ Awareness on latest AI technologies such as Github...PlatformCloudLocal areaWorldwide
$118.3k - $219.8k
Are you excited to lead Site Reliability Engineering teams that keep mission-critical, 24/7 services... ...securely?Do you enjoy building automated cloud platforms, hardening security, and driving... ...About our TeamOur globally distributed SRE team operates across the FCC market,...PlatformCloudFull timeLocal area- ...Job Description: AI/ML Engineer (2026) Role Summary: We are looking for a highly... ...focus on turning experimental models into reliable, high-performance, real-world... ...resilient, and secure AI infrastructure on cloud platforms (AWS, Azure, or GCP). Data Engineering...PlatformCloud
- ...Salesforce Developer - Agentforce (Salesforce AI-driven solutions and automation) -... ...with strong expertise in the Salesforce platform and hands-on experience with Agentforce (... ...implement solutions on Salesforce (Sales Cloud, Service Cloud, Experience Cloud) · Build...PlatformCloudTemporary work
$120k - $160k
...generations to come. Job Purpose and Impact ~ The AI Security Engineering Manager will help solidify foundation for the company's... ...business applications Collaborate with AI platform engineers, cloud engineers, software engineers and other key stakeholders...PlatformCloud- ...Job Title AI/ML Engineer Job Summary We are looking for a talented AI/ML Engineer to design, develop, and deploy machine... ...with Docker and Kubernetes Experience with cloud platforms: Amazon Web Services Google Cloud Microsoft...PlatformCloud
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Platform & Cloud. Be the first to apply!
- site reliability engineer Alpharetta, GA
- site reliability engineer sre Alpharetta, GA
- platform developer Alpharetta, GA
- platform engineer Alpharetta, GA
- senior aws cloud engineer Alpharetta, GA
- aws cloud architect Alpharetta, GA
- informatica cloud developer Alpharetta, GA
- senior principal cloud computing engineer Alpharetta, GA
- cloud network engineer Alpharetta, GA
- cloud engineer Alpharetta, GA



