Lead Site Reliability Engineer
$125k - $175kMorgan Stanley
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.
Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals.
Interested in joining a team that’s eager to create, innovate and make an impact on the world? Read on.
The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.
This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues.
What you’ll do in the role:
Drive Reliability Engineering Practices
Champion SRE principles, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational risk management frameworks.
Define and track reliability metrics that measure platform health, customer experience, and operational effectiveness.
Lead initiatives focused on reducing Mean Time to Detect (MTTD), Mean Time to Identify (MTTI), and Mean Time to Restore (MTTR).
Own Production Reliability
Provide leadership for the proactive detection, triage, and resolution of production issues impacting business-critical applications and services.
Serve as the primary owner for escalated production incidents, driving resolution efforts across application, infrastructure, vendor, and external partner teams until service is restored and client impact is mitigated.
Establish a culture of operational excellence focused on stability, resiliency, and continuous service improvement.
Ensure clear, concise, and timely communication during outages, providing accurate business impact assessments and recovery updates to senior leadership.
Production Governance & Change Management
Serve as a key gatekeeper for the production environment, ensuring adherence to change management policies, release controls, operational readiness standards, and risk management practices.
Assess the operational impact of technology changes and ensure appropriate testing, rollback strategies, monitoring, and support models are in place prior to production deployment.
Partner with development teams throughout the software lifecycle to ensure reliability, observability, and operational supportability are built into new applications and services.
Automation & Operational Efficiency
Identify opportunities to eliminate manual effort and operational toil through automation, self-healing capabilities, and AI-driven operational workflows.
Lead the development and adoption of automation solutions that improve reliability, reduce risk, and increase operational efficiency across the organization.
Promote a culture of engineering-led operations and continuous process optimization.
Operational Readiness & Knowledge Management
Establish and maintain a comprehensive knowledge management framework, ensuring runbooks, troubleshooting guides, standards, and operational procedures are accurate, current, and accessible.
Drive operational readiness programs that improve first-level diagnosis and reduce dependency on development teams for routine issue resolution. Create End-to-End Know your system diagrams.
Ensure support teams maintain high-quality documentation and standardized troubleshooting practices to accelerate incident resolution.
Technical Leadership
Act as a senior technical leader and trusted advisor for reliability, resiliency, observability, and production support strategies.
Provide guidance on architecture reviews, platform scalability, capacity planning, disaster recovery, and resiliency testing initiatives.
Partner with engineering teams to identify systemic risks and implement long-term solutions that improve platform stability and customer experience.
What you’ll bring to the role:
10+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end-to-end troubleshooting, networking fundamentals and appropriate attention to detail
BS/MS or equivalent, preferably in quantitative discipline (Computer Science, Computer Engineering).
5+ years’ experience in leading a small to medium team of alike skillset.
5+ years of experience in driving SRE principles and Chaos Engineering.
Experienced, technically hands-on professional that understands both code and infrastructure
Strong experience in scripting language (Shell scripting, Python, Perl, etc.) and cloud driven development
Strong database skills with DB2, Sybase or Oracle
Hands-on experience with Autosys or other batch scheduling software
Experience in AWS/GCP/Azure Cloud technologies
Working knowledge on any of the DevOps & observability tools (Grafana, Prometheus, Splunk, Kibana)
Solid analytical skills, problem determination, and resolution recovery processes
Ability to interface and cultivate excellent working relationships with technology teams, business analysts, and vendors
Experience in web analytics tools (preferably Adobe Experience Cloud tools) is Plus
Should be a fast learner of technologies in a quick paced environment.
Have strong organizational skills and the ability to manage multiple tasks and high-pressure situations for outage handling, management, or resolution
Is driven to learn about new technologies, techniques and what it takes to be an integral member of this team
Hands-on experience administering large-scale, high-availability systems and the tools to monitor performance and availability
Excellent communication and writing skills specific to technical discussions across the management layers
Experience with incident “on call” and ability to respond to emergencies on a 24/7 basis
Hands-on with AI and implementation of AI tools for operational efficiency
Strong ownership mentality with a focus on customer satisfaction
Be able to manage an outage incident, coordinating user communications, and other teams to help resolve an incident.
Experience working with Financial Services area will be a plus
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Expected base pay rates for the role will be between $125,000 and $175,000 per year at the commencement of employment. However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit: .
- ...ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability... ...issues, identify root causes, and implement permanent fixes.Lead post-incident reviews, create action items, and follow...SuggestedPermanent employmentFull timeWork experience placementLocal area
$86.6k - $144.4k
...automation to solve complex security and reliability challenges?Do you enjoy shaping the... ...LexisNexis Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role in ensuring... ...Site Reliability Engineer (SRE) to lead and scale security, compliance, and...SuggestedFull timeLocal area- ...and shape the future of our communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job... ...the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of...SuggestedFlexible hoursWeekend work
$128k - $216k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe your... ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is... ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability...SuggestedFull timeWorldwide$125k - $175k
...power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment...SuggestedTemporary work- ...of our communities.This is a Software Engineering position at Director level, which is part... ...role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform... ...incident responseDemonstrated ability to lead RCAs, deploy fixes, and drive...
$167.7k - $245.2k
...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze...Full timeTemporary workLocal areaFlexible hours$118.3k - $219.8k
Are you excited to lead Site Reliability Engineering teams that keep mission-critical, 24/7 services running reliably and securely?Do you enjoy building automated cloud platforms, hardening security, and driving ongoing cost optimization through strong FinOps practices?...Full timeLocal area- ...HybridAt PDI Technologies, we empower some of the world's leading convenience retail and petroleum brands with cutting-edge... ...OverviewPDI Technologies is looking for a Senior Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments...Full time
- A technology solutions company is seeking a Senior Databricks Administrator located in Georgia. This role involves designing and optimizing a scalable Databricks platform for AI and ML workloads. The ideal candidate has proven experience with Terraform, strong programming...Contract work
- ...On-Site role Job Description: ~4 - 10 hour days (Sunday-Wednesday 7am-5pm) ~ Database Site Reliability Engineer (Database Operations) Position Summary: We are seeking... ...incidents and service degradations. Lead troubleshooting efforts for...Permanent employmentTemporary work
- ...are empowered to shape the future of payments.As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE... ...Our proven, secure and scalable software solutions enable leading corporations, fintechs and financial disruptors to process...Full timeLocal areaWorldwide
$48 per hour
...location is currently looking for a Lead Heavy Duty Mechanic to join... ...to ensure equipment reliability and performance. Complete and... ...with the various types of diesel engines; ~ Ability to lead in a team... ...environment, Clean Harbors is on-site, providing premier...Full timeCasual workWeekend work- ..., working closely with Dev team to improvise application reliability and stability.Preferred Skill and Experience• Experience... ...Management|ITIL Technical/Domain Skill 4Technology|DevOps|Site Reliability Engineering(SRE) Work LocationAlpharetta, GA, New York, NYCountryUSAState...Full timeTemporary workRelocationShift work
- ...Functional Lead + SAP FI Consultant Location: Alpharetta, GA Duration: Fulltime Job Description: Skills Desired: 1. SAP... ...FI. Lead MTO (Make-to-Order), MTS (Make-to-Stock), and ETO (Engineer-to-Order) cost object controlling production order costing, WIP...Full timeImmediate start
- ...Overview: About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts... ...thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability...Full timeLive inWork at office
$129k - $161k
...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department... ...0 About Priority Commerce: Priority Commerce is a leading financial technology company on a mission to deliver a...Remote work- Company DescriptionAmerica Networks is a leading sensor and networking solutions partner for companies in any Industrial, Manufacturing... ...asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing,...Work experience placement
$94.1k - $141.1k
...assigned office location, if available.Our Lead Strategic Accounts 3 CNV, earn between $9... ...engaging in sales activities at customer site; communicating with customers via phone,... ...AT&T external partners, including design/engineering; researching customer business/industry...Temporary workWork at officeLocal area- ...FLSA Status: ExemptJOB SUMMARYThe Security Lead is responsible for the City's information... ...remote access services for staff, remote sites, and third-party vendors, including... ..., Information Technology, Cybersecurity, Engineering, or a closely related field required.At least...Contract workCasual workLocal areaRemote work
- ...committed to excellence.POSITION OVERVIEWThe Enterprise AI Enablement Lead is a senior individual contributor responsible for designing,... ...AI adoption at scale.ESSENTIAL RESPONSIBILITIESAI Solution Engineering & ArchitectureIdentify and translate ambiguous business...Permanent employmentFull timeTemporary workWork experience placementWork at office
$125k - $175k
...our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Application Security Engineering position at the Vice President level, which is part of the job family responsible for providing specialist data...Temporary work$24.05 - $31.25 per hour
...changing demands. We’re committed to providing sustainable and reliable solutions to our customers. Join our team to help build a better... ...safety, quality, and continuous improvement. We are seeking a Team Lead who will support daily operations, drive team performance, and...Full timePart timeRelocation packageFlexible hoursShift work$50k - $80k
...are powered by extraordinary people. Our innovative products and reliable services are delivered with convenience, excellence, and quality... ...'s Committee. Position Summary: The Field Canvassing Team Lead is responsible for hiring, training, and developing a team of Field...Full timeH1bWork at officeLocal areaWork from homeShift workAfternoon shift- ...independent judgment to assure successful and safe delivery of care. Facilitates onboarding and precepting under the guidance of their direct lead. Serves as a clinical resource/support to the staff, physicians, patients, families, and other departments by providing direct...Full timeShift workNight shift
- ...Lead Structured Cabling TechnicianThe Lead Structured Cabling Technician is responsible... ...with the project manager, network engineers, and site supervisors to report progress and... ...quality inspections to guarantee system reliability and performance.Maintain detailed records...For contractorsLocal area
- ...to be empowered to grow? Hiring immediately for part-time Shift Leads - we're ready for you! BENEFITS: Daily pay - work... ...of Cashier experience Must have valid driver's license and reliable transportation Must be able to perform repeated bending, standing...Daily paidPart timeLocal areaImmediate startShift work
$120.2k - $201.8k
...engaged as the technical primary project lead and architect. This seasoned consultant... ...ethics, trustworthiness, accountability and reliability in being resourceful and self-directed... ...selection, sizing and detailed engineered designData center expertise including core...Temporary workWork at officeLocal areaRelocation1 day per week- ...and subcontractors; and overseeing day-to-day project execution.Lead and/or participate in initial and ongoing project planning meetings... ...policies, values, and business practices. Conduct project site inspections to monitor progress, support project teams, and proactively...Temporary workFor contractorsFor subcontractor
$105.05k - $161.8k
...complex statistical and financial models for forecasting and defines KPIs and metrics to measure business performance.Responsibilities• Leads complex data and business analysis to develop business plans, and identifies recommendations and insights.• Works independently to...Full timeTemporary workWork experience placementLocal areaFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead operating engineer Alpharetta, GA
- lead engineer Alpharetta, GA
- site reliability engineer Alpharetta, GA
- site reliability engineer sre Alpharetta, GA
- on-site clinical research associate (traveling/remote) Alpharetta, GA
- site safety Alpharetta, GA
- junior website developer Alpharetta, GA
- construction site safety Alpharetta, GA
- IT site lead Alpharetta, GA
- website content developer Alpharetta, GA





