Site Reliability Engineer (SRE) - AI Platform & Cloud
Morgan Stanley
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.
This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.
Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
What you'll do in the role:
Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
Design and build automation for core platform capabilities, reducing manual toil
Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
Optimize cost vs. performance tradeoffs in large-scale compute environments
Harden systems for security, compliance, auditability, and data governance
Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
Maintain runbooks, operational playbooks, documentation, and training materials
Participate in on-call rotations and respond to production incidents 24/7 as needed
Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability
What you ' ll bring to the role:
Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience
5 years of production experience in SRE / Infrastructure / ops for large-scale systems
Strong programming/scripting skills (Python, Go, Java, or equivalent)
Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
Solid experience in capacity planning, performance tuning, scaling, and incident response
Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
Excellent communication, documentation, and cross-team collaboration skills
Proven track record of reducing operational toil via automation
Nice to have
Understanding of SRE techniques.
Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
Understanding of ModelOps/ ML Ops/ LLM Op.
Experience with chaos engineering, canary deployments, blue/green rollouts
We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.
All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren’t just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.
Wherever you are in our 1,200 global offices, you’ll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.
At Morgan Stanley Alpharetta, we support the Firm’s global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.
Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.
It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.
Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit: .
$60 - $65 per hour
...Role Overview The Site Reliability Engineer will support Cyber Data Risk & Resilience... ...of critical cybersecurity platforms and services. This role is... ..., APIs, data pipelines, and cloud components to provide end-to... ...Experience with cloud-based AI services such as Azure AI,...PlatformCloudContract work- ...Cloud Security Engineer – SRE We are seeking a skilled and motivated Cloud Security Engineer - SRE... ...focus on solution engineering/site reliability. This role will involve collaborating... ...Cloud Computing: Knowledge of cloud platforms such as AWS, Azure, or Google Cloud...PlatformCloud
- ...Overview: Job Title: AI/ML Ops & Infrastructure Engineer Company: R2... ...AI & IoT Intelligence Platform utilizing advanced NLP... ...Kubernetes across multi-cloud environments (AWS, GCP,... ...experience in MLOps, DevOps, Site Reliability Engineering (SRE), or Cloud...PlatformCloudFull timeRemote workShift work
- ...SRE (Security Engineer) We are seeking a skilled and motivated Cloud Security Engineer - SRE to join our dynamic team. The ideal... ...on solution engineering/site reliability. This role will involve collaborating... ...: Knowledge of cloud platforms such as AWS, Azure, or...PlatformCloudWork experience placement
- ...Job Posting Title Cloud Developer Location... ...on Generative AI you will design and... ...agentoriented software engineering AOSE journey You... ...engagement platform leveraging the latest... ...productionready and meet reliability standards... ...of experience in SRE DevOps MLOps infrastructure...PlatformCloud
$100k - $150k
...to grow, we’re looking for a skilled Site Reliability Engineer (SRE) to join our dynamic team and contribute... ..., and continually pushing the platform toward higher reliability with lower... ...on experience with at least one major cloud platform (AWS, Azure, or GCP). Background...PlatformCloudFull timeH1bLocal areaImmediate startRemote workVisa sponsorshipWork visa- ...Job Title: GCP Cloud Engineer with SRE Location: Alpharetta GA - Day 1 Onsite Duration: 6 to 12 Months 1 Cloud Engineer with SRE experience... ...: Cloud Engineer will be part of the GCP Cloud Platform team who is responsible for building Cloud Infrastructure...PlatformCloud
- ...security sector, is seeking a Security Engineer IV (Cloud Security Engineer - SRE) to join their innovative team.... ...the Cloud Security and Site Reliability Engineering teams. The ideal candidate... ...configurations. Experience with cloud platforms like AWS, Azure, or Google Cloud,...PlatformCloud
- ...Application Support Engineer Location :... ...Production Management Site Reliability Engineer position... ...Management application platforms, participation on... ...perform DevOps/SRE role in Java, Unix... ...multi-tiered cloud-based applications... ...related field AI tools may assist in...PlatformCloudHourly payPermanent employmentContract work
- ...Senior AI Engineer Equifax is where you can power your possible... ...recommend AI frameworks, cloud services, and third-party platforms aligned with business... ...engineering, quality engineering, reliability engineering and project... ...Influence architects, SRE leads and other technical...PlatformCloudFull timeWork at officeImmediate startRemote workMonday to Friday
- ...Machine Learning Engineer / Data Scientist We are seeking a skilled... ...with expertise in Google Cloud Platform (GCP), Databricks, Kubernetes, and Generative AI. This role involves designing... ...systems to ensure performance and reliability in production environments....PlatformCloud
- ...Overview: Job Title: Data Engineer (AI & Data Platforms) Company: R2 Technologies Location... ...downstream analytics. Implement data reliability engineering practices (data contracts... ...SQL. Hands-on experience with cloud data platforms (Snowflake, Databricks...PlatformCloudFull timeRemote workShift work
$142.6k - $261.5k
...opportunity The Platforms Practice... ...our product-driven, AI-centric approach,... ...designers, and software engineers enable our clients... ...building and operating cloud infrastructure and... ...with a focus on reliability and excellent... ...across teams Apply SRE best practices, establish...PlatformCloudSummer holidayFlexible hours- ...Java, data warehousing, cloud computing, and build engineering. As a global provider of... ...We are seeking a skilled AI & Enterprise Applications... ...high system performance, reliability, and scalability ~... ...Experience with enterprise platforms and system integrations...PlatformCloudShift work
- ...Salesforce Developer w/ AI Brokerage Location : Alpharetta... ..., GA -hybrid 3 days on site Duration 6 -18mth+ Contract... ...prompt templates Data Cloud capabilities like data models,... ...Proficient to a cloud computing platform and the associated automation...PlatformCloudContract workShift work
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...innovation! Java Backend Engineer (AI & ML) Location:... ...secure, performant, and reliable while handling AI/ML workloads...PlatformCloudFull time
- ...specializing in Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure), Artificial Intelligence (AI), Machine Learning (ML),... ...1000 companies-with scalable, platform-based solutions and data-driven... ...in Computer Science, Software Engineering, or a related field (or equivalent...PlatformCloudFull time
- ...specializing in Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure), Artificial Intelligence (AI), Machine Learning (ML),... ...1000 companies-with scalable, platform-based solutions and data-driven... ...in Computer Science, Software Engineering, or a related field (or equivalent...PlatformCloudFull time
- ...Role: ML/AI Engineers (This role is open to US Citizens, Green Card holders, GC-EAD only. We do not sponsor... ...and developing AI or machine learning solutions on platforms such as AWS, Databricks, Azure, Google Cloud and OpenAI. Software engineering and/or Data Engineering...PlatformCloudRemote workVisa sponsorshipRelocation package
- ...specializing in Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure), Artificial Intelligence (AI), Machine Learning (ML),... ...1000 companies-with scalable, platform-based solutions and data-driven... ...in Computer Science, Software Engineering, or a related field (or equivalent...PlatformCloudFull time
- ...Machine Learning, Java, data warehousing, cloud computing, and build engineering. As a global provider of IT staffing... ...NoSQL) Familiarity with cloud platforms and containerization tools... ...event-driven systems Exposure to AI or modern automation technologies...PlatformCloudShift work
- ...Exposure to data preprocessing tools (Pandas, NumPy) - Basic cloud AI services Mid Level (5–12 years) Mandatory Skills: -... ...Skills: - Experience with NLP, deep learning - Cloud ML platforms (AWS SageMaker, Azure AI) Expert (12+ years) Mandatory...PlatformCloud
- ...AI Software Engineer III We are CirrusLabs. Our vision is to become the world's most sought-after niche digital transformation company... ...ML architectures, tools, and data infrastructure to public cloud platforms. Expert level programming skills in two or more...PlatformCloud
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML... ...companies-with scalable, platform-based solutions and data... ...for scalability, reliability, and security in cloud... ...Computer Science, Software Engineering, or a related field (...PlatformCloudFull time
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...! Angular DevSecOps Engineer (CI/CD, Kubernetes) Location... ...with Kubernetes for reliable production environments...PlatformCloudFull time
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...innovation! Python Data Engineer (Airflow, Kafka, Spark)... ...pipelines for performance, reliability, and scalability in...PlatformCloudFull time
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...Senior Python AI/ML Engineer (TensorFlow/PyTorch)... ...to maintain production reliability. Required Qualifications...PlatformCloudFull time
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...innovation! Cloud Data Engineer (GCP BigQuery & Dataflow... ...analysts to provide clean, reliable data for reporting and...PlatformCloudFull time
- ...Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure... ...Intelligence (AI), Machine Learning (ML)... ...companies-with scalable, platform-based solutions and data... ...innovation! Python Backend Engineer (FastAPI & Docker)... ...performance, security, and reliability in production....PlatformCloudFull time
- ...specializing in Java, .NET, Big Data, Cloud Computing (AWS, GCP, Azure), Artificial Intelligence (AI), Machine Learning (ML),... ...1000 companies-with scalable, platform-based solutions and data-driven... ...in Computer Science, Software Engineering, or a related field (or equivalent...PlatformCloudFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Platform & Cloud. Be the first to apply!
- client platform engineer Alpharetta, GA
- platform developer Alpharetta, GA
- platform engineer Alpharetta, GA
- aws cloud infrastructure engineer Alpharetta, GA
- remote cloud architect Alpharetta, GA
- senior cloud engineer Alpharetta, GA
- cloud architect Alpharetta, GA
- cloud engineer remote Alpharetta, GA
- senior principal cloud computing engineer Alpharetta, GA
- software engineer - cloud services Alpharetta, GA

