Site Reliability Engineer (SRE) - AI Platform & Cloud
Morgan Stanley
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.
This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.
Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
What you'll do in the role:
Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
Design and build automation for core platform capabilities, reducing manual toil
Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
Optimize cost vs. performance tradeoffs in large-scale compute environments
Harden systems for security, compliance, auditability, and data governance
Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
Maintain runbooks, operational playbooks, documentation, and training materials
Participate in on-call rotations and respond to production incidents 24/7 as needed
Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
What you ' ll bring to the role:
Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
5 years of production experience in SRE / Infrastructure / ops for large-scale systems
Strong programming/scripting skills (Python, Go, Java, or equivalent)
Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
Solid experience in capacity planning, performance tuning, scaling, and incident response
Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
Excellent communication, documentation, and cross-team collaboration skills
Proven track record of reducing operational toil via automation
Nice to have
Understanding of SRE techniques.
Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
Understanding of ModelOps/ ML Ops/ LLM Op.
Experience with chaos engineering, canary deployments, blue/green rollouts
We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.
All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren’t just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.
Wherever you are in our 1,200 global offices, you’ll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.
At Morgan Stanley Alpharetta, we support the Firm’s global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.
Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.
It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.
Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit: .
- ...We are looking for a SRE/DevOps Engineer A highly technical... ...enterprise infrastructure and platform modernization... ...DevOps practices and AI-enabled tooling.... ...delivery through DevOps and Site Reliability Engineering best... ...Assist with AWS cloud infrastructure deployments...PlatformCloud
- ...role involves designing and optimizing a scalable Databricks platform for AI and ML workloads. The ideal candidate has proven experience with... ...and Java, and hands-on experience working with Databricks and cloud services. This is a contract position oriented towards...PlatformCloudContract work
- ...leading workflow management platform built exclusively for K-12... ...the future of education. Site Reliability Engineer (SRE) Overview: We are... ...just operate a dashboard. AI-Accelerated Execution (core... ...Automation, Infrastructure & Cloud: Proficient in Python, Go,...PlatformCloudFull timeLive inWork at office
- ...have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite).... ...2. Terraform 3. SRE 4. Python 5... ...Maven 6. GCP or any Cloud Job Requirements (7+ years... ...• Experience with Cloud Platforms and virtualization Technologies...PlatformCloudImmediate startRelocation
- ...Production Management & Reliability Engineering position at Director... ...organization delivers platforms that support core client... ...We are looking for a Site Reliability Engineer... ...Python, Perl, etc.) and cloud driven development ~... ...leveraging generative AI tools to enhance...PlatformCloudFull timeFlexible hoursWeekend work
- ...scripting is a plus Provisioning cloud infrastructure (AWS, GCP) using... ...Collaborating with product, architecture, and engineering groups to build a platform that streamlines application... ...level experience in DevOps, Site Reliability Engineering with expertise in Enterprise...PlatformCloud
- ...Role: ML/AI Engineers (This role is open to US Citizens, Green Card holders, GC-EAD only. We do not sponsor... ...and developing AI or machine learning solutions on platforms such as AWS, Databricks, Azure, Google Cloud and OpenAI. Software engineering and/or Data Engineering...PlatformCloudRemote jobFull timeVisa sponsorshipRelocation package
- ...forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of... ...: 6+ years as an SRE, DevOps Engineer, or similar role Cloud: Strong experience with AWS (EKS, EC2, S3, Route53, IAM)...CloudPermanent employmentWork experience placementLocal area
$129k - $161k
...Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20... ...engineering experience, including 3+ years in SRE or reliability-focused roles. ~ Demonstrated...PlatformCloudRemote work- ...We are looking to add a Software Engineer (Level II) to our Development team to help build... ...exposed to many modern technology platforms and cloud-based applications in the market. You... ...Ability to demonstrate effective use of AI coding tools like Claude code ~ Reside...PlatformCloudFull timeWork at officeRelocation
- ...relocate on own Job Title: Cloud Solutions Architect Duration... ...modern cloud-based data and AI platforms. The ideal candidate will... ...including AI, GenAI, ML, and data engineering innovations. Define... ...platform performance, scalability, reliability, and cost efficiency....PlatformCloudLocal areaImmediate startRelocation
- ...Doeren Mayhew is seeking a Senior Software Engineer. The role may be based in Troy, Michigan... .... Identify opportunities to introduce AI‑driven capabilities, automation, and... ...design and SQL optimization. Exposure to cloud platforms (Azure, AWS, or GCP) and modern deployment...PlatformCloud
$120k - $160k
...generations to come. Job Purpose and Impact ~ The AI Security Engineering Manager will help solidify foundation for the company's... ...business applications Collaborate with AI platform engineers, cloud engineers, software engineers and other key stakeholders...PlatformCloud- ...classification, data extraction) by orchestrating OCR/AI services from Pega workflows and managing... ...DB Trace, and PDC (Predictive Diagnostic Cloud); tune data pages, query strategies,... ...rates. Conduct impact assessments for platform upgrades (Pega version/hotfixes), ruleset...PlatformCloud
- ...innovative Intermediate Data Engineer to join our Global Keying & Linking... ...engineer or related role. Cloud certification strongly preferred... ...experience with Google Cloud Platforms and an overall understanding... ...building, or maintaining agentic AI workflows and solutions using...PlatformCloudWork at officeRemote workMonday to Friday
- ...is hiring a Senior Data Engineer to help build and modernize the lottery data platform that supports reporting,... ...data science, and future AI capabilities. You will build and operate reliable data pipelines, improve... ...environments across legacy and cloud systems. What This...PlatformCloudLocal area
- ...Oracle Services Apps Tech Manager to lead enterprise reporting and AI-driven analytics projects. You will guide senior client stakeholders, define reporting strategies, and design cloud-based data platforms across ERP, HCM, SCM, and CX. You will manage distributed teams,...PlatformCloud
- ...solutions. This role blends hands‑on technical leadership with architectural governance, AI‑first solution design, and cross‑platform integration. The ideal candidate brings deep Service Cloud expertise, strong integration and security knowledge, and the maturity to mentor...PlatformCloudContract workTemporary work
$200k - $250k
...Employer - Synthetix Labs – AI Products Group at Celsior... ...vision, product engineering strategy, AI architecture, platform scalability, security, and... ...traceability, security, and reliability. The product has been incubated... ...AI, digital engineering, cloud transformation,...PlatformCloudImmediate start- Equifax is seeking a driven Senior Data Engineer to contribute to the creation of high-... ...highly visible and critical Ignite AI Advisor Application. You will partner... ...with large datasets on a big data platform (e.g., Google Cloud, AWS, Snowflake, Hadoop). Intermediate...PlatformCloudWork at officeRemote workMonday to Friday
- ...seeking a highly skilled and forward-thinking AI Engineer to join our AI Engineering team. This... ...and decision-making across our FinTech platforms. This is a unique opportunity to work at... ...workflows using modern frameworks and cloud-native tools. Collaborate with data scientists...PlatformCloudH1bWorldwide
$92.16k - $138.24k
...years of experience working with Google Cloud Platform Preferred Skills: · 10+... ...innovation. We are one of the world's leading AI and digital infrastructure providers,... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and...PlatformCloudTemporary workWork at officeRemote workFlexible hours$76.2k - $174.1k
...Manager – Financial Services Organization – AI and Data Service Delivery Center EY is... ...of related work experience in AI/ML engineering or MLE/ML Ops Experience working with... ...PyTorch, etc.) Extensive experience with cloud platforms such as AWS, Azure, or Google Cloud...PlatformCloudWork experience placementSummer holidayFlexible hours$123.62k - $267.75k
...responsible for developing the vision for AI-enabled data ecosystems, defining the... ...Hands-on experience with modern data and AI platforms such as Snowflake, Databricks, Azure, GCP... ...the supporting infrastructure, including cloud and cybersecurity technologies, required...PlatformCloudLocal area- ...AI Engineer Equifax is where you can power your possible. If you want... ...architecting and deploying cutting-edge, cloud-native solutions for a large... ...and build highly scalable, reliable, and performant APIs, microservices, and PaaS/SaaS platforms, including the ability to...PlatformCloudFull timeWork at officeImmediate start3 days per week
- ...Data Scientist with Python and AI/ML consultant. Location:... ...analytics across enterprise platforms. The role focuses on transforming... ...Perform data exploration, feature engineering, and model evaluation. Build... ...deploying models in cloud environments (AWS/Azure/GCP)....PlatformCloudContract work
- ...seeking an experienced Senior AI Developer to design, develop,... ...), backend development, and cloud-based AI services. The role involves... ...solutions into enterprise platforms. Develop APIs, microservices... ..., embeddings, prompt engineering, and RAG architectures. Collaborate...PlatformCloudFull timeRemote work
- ...The Technical Operations Engineer is responsible for supporting the performance, reliability, and visibility of the Medlytix Production... ...monitoring tools, telemetry platforms, cloud technologies, and data... ...actionable insights Exposure to ML/AI concepts, tools, or...PlatformCloud
$110k - $186k
...times a day - quickly, reliably, and securely. Any... ...Senior DevOps Engineer About your role: As... ...Client Implementations platforms. You will... ...automation solutions using cloud native services and... ...tools incorporating AI into daily... ...role: This role is on-site Monday through Friday...PlatformCloudFull timeContract workTemporary workFor contractorsH1bWork at officeLocal areaMonday to Friday$101k - $194k
...state-of-the-art agentic AI/ML models and systems... ...code while ensuring reliability, performance, and scale... ...systems for VCG platforms. Integrating AI agent... ...product managers and GTS engineering teams to implement and... ...development. Experience with cloud platforms (e.g., GCP,...PlatformCloudFull timeTemporary workPart timeWork experience placementWork at officeWork from homeShift work3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Platform & Cloud. Be the first to apply!
- cloud developer Alpharetta, GA
- senior cloud data engineer Alpharetta, GA
- senior aws cloud engineer Alpharetta, GA
- cloud engineer Alpharetta, GA
- senior cloud network engineer Alpharetta, GA
- informatica cloud developer Alpharetta, GA
- cloud architect Alpharetta, GA
- site services specialist Alpharetta, GA
- construction site safety Alpharetta, GA
- site leader Alpharetta, GA




