Site Reliability Engineer
Morgan Stanley
Overview In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This Lead Software Production Management & Reliability Engineering position is at Director level and is responsible for overseeing the production environment, ensuring operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime. Job Summary We are looking for a Site Reliability Engineer with a minimum of 5 years of industry experience, preferably in the financial IT community. The role focuses on production support within the WM Product Technology team, automating deployments and working with agile teams to build and support stable and reliable production systems. The ideal candidate will be passionate about automation and proficient in programming languages such as Python, PERL, SHELL, Ruby, Java, or C# and have a strong understanding of database concepts, job schedulers (e.g., Autosys), MQ, web services, UNIX/Linux/Windows OS, and debugging applications. The candidate should also be a strong leader with excellent communication skills, organized, disciplined, detail‑oriented, self‑motivated, and delivery‑focused. Responsibilities Maintain applications after deployment by measuring and monitoring availability, latency, and overall system health, focusing on business activities and continuously evaluating cost and TOIL. Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation, capacity planning, and launch reviews. Scale systems sustainably through automation and evolve them by pursuing changes that improve reliability and velocity, including operational automation. Troubleshoot infrastructure issues, review log files, update documentation, and maintain a knowledge base with resolutions. Collaborate closely with the application development team to understand the platform and create tools/utilities that aid production management. Work with upstream data providers and consumers to reduce escalation to development teams. Develop scripts and assist with code changes along with operational tasks and activities. Ensure the support team has excellent knowledge of the application set, owns and maintains the support knowledge base and documents. Use analytical skills to identify trends in the environment and drive problem resolution. Lead efforts to determine improvement areas to stabilize the production environment. Identify risks and act with urgency, working within a team or independently. Test and tune network, hardware, and software configurations to maximize performance. Interface with teams such as IT Dev managers and infrastructure teams and serve as a Subject Matter Expert (SME) for supported applications. Take ownership of and manage production requests, questions, issues, and perform root cause analysis for outages and incidents. Be flexible to provide weekend on‑call rotation and be available for offshore time lead. Be accountable for the Production and non‑Production environments and be part of 24/7 production support coverage. Skills Required 5+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end‑to‑end troubleshooting, networking fundamentals, and attention to detail. Ability to focus and provide resolutions for production issues in a high‑demand, pressured environment. 5+ years of hands‑on experience designing, developing, and implementing technical solutions, or significant experience in deep technical support. Strong experience in scripting languages (Shell scripting, Python, Perl, etc.) and cloud‐driven development. Strong database skills with DB2, Sybase, or Oracle. Hands‑on experience with Autosys or other batch scheduling software. Strong experience in continuous integration and continuous deployment. Strong experience with on‑demand environments for both virtual machines and containers. Knowledge and hands‑on experience with monitoring tools such as Splunk, IP Soft, Sockeye. Practical experience in Agile methodology (e.g., Scrum). Knowledge or experience with automating deployments using Jenkins and Train. Ability to diagnose technical problems, debug, optimize code, and automate routine tasks. Hands‑on experience in application and database troubleshooting/issue resolution in a fast‑paced environment. Excellent communication and ability to think creatively for process improvements. Knowledge of cloud‑based deployment, security, networking concepts in Azure and AWS. Hands‑on experience leveraging generative AI tools to enhance research, automate, and improve productivity. Skills Desired Knowledge or experience with algorithms, data structures, complexity analysis, and software design. Interest in designing, analyzing, and troubleshooting large‑scale distributed systems. Educational Qualification Minimum BS degree in Computer Science, Engineering, or a related field. Equal Employment Opportunity Statement It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law. Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet). #J-18808-Ljbffr
- 4+ years of experience in an SRE, DevOps, or cloud infrastructure role. Strong experience with Azure cloud services and infrastructure. Hands-on experience with java and Terraform and Terragrunt for infrastructure-as-code. Proficiency with Kubernetes (preferably AKS), Databricks...Suggested
- ...Overview Location: Alpharetta, GA (3 days a week onsite) Duration: 6 months Job Description: We are seeking a skilled Site Reliability Engineer to join our team and help build, maintain, and scale our cloud-native infrastructure. You will work closely with development...Suggested3 days per week
$60 - $63 per hour
...learn more. Base pay range $60.00/hr - $63.00/hr Our client, a global information and analytics company, is looking for a Site Reliability Engineer to join their team in Alpharetta, GA! This is a 12-month initial contract and is hybrid so local candidates are required....SuggestedFull timeContract workTemporary workLocal areaFlexible hours- ...Job Title: Site Reliability Engineer (SRE) Experience Required: 8+ Years Industry: Banking / Financial Services Job Description We are seeking an experienced Site Reliability Engineer (SRE) to join our banking client’s technology team. The ideal candidate...Suggested
- ...I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated resume or if you could refer someone I would really appreciate it. Role...SuggestedImmediate startRelocation
- ...initiatives. Collaborating with product, architecture, and engineering groups to build a platform that streamlines application... ...Experience: 5+ Years advanced level experience in DevOps, Site Reliability Engineering with expertise in Enterprise Cloud infrastructure...
- ...Overview Are you an experienced Site Reliability Engineering leader ready to shape strategy, inspire teams, and drive innovation at scale? Are you looking to lead a high-impact SRE team where your leadership will directly influence innovation, reliability, and engineering...Full timeRemote workRelocation
$129k - $161k
...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20 About Priority: Priority Technology Holdings, Inc. is a leading...Remote work- ...security to responsibly propel the global lottery industry ever forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with...Permanent employmentWork experience placementLocal area
- ...communities. This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family... ...the business. Job Summary We are looking for a Site Reliability Engineer with a minimum of 5 years of industry...Full timeFlexible hoursWeekend work
- ...We are looking for a SRE/DevOps Engineer A highly technical, hands-on engineer to join a critical engineering organization supporting... ...Python. Improve software delivery through DevOps and Site Reliability Engineering best practices. Support Infrastructure as...
- ...Job Title Google Cloud Site Reliability Engineer (GCP SRE) - AI/Vertex AI Job Summary We are seeking an experienced Google Cloud Site Reliability Engineer (GCP SRE) with expertise in Google Cloud Platform (GCP) , Site Reliability Engineering , Kubernetes...
- ...shape the future of our communities. This is a Software Engineering position at Director level, which is part of the job family... ...businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support...Full time
- ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid – large wireless...
$63.22k - $94.55k
...maintain Windows Server environments with a focus on performance, reliability, and scalability Develop and maintain PowerShell scripts... ...excellence and process optimization Mentor junior engineers and contribute to knowledge sharing across the team Required...Remote work$63.22k - $94.55k
...States (US). The role involves designing, implementing, and maintaining Windows Server environments with a focus on performance, reliability, and scalability. The Lead Systems Programmer will also develop and maintain PowerShell scripts for automation of administrative...Remote work$115.4k - $192.3k
...seeking a highly experienced Consulting / Principal Software Engineer to lead the design, optimization, and management of large-scale... ..., CTEs, and window functions to improve performance and reliability Perform hands-on database programming by developing and maintaining...Temporary workLocal area- ...The Lead Operating Engineer is responsible for the HVAC system and all mechanical equipment within the building. The position works very closely with the Chief Engineer to ensure that the building systems are functioning properly. Primary Functions Monitor...For contractorsWork at office
$28.9 - $34 per hour
...Job Title Mobile Building Engineer Job Description Summary The position involves maintaining and repairing HVAC, plumbing, electrical, and building mechanical systems to ensure maximum efficiency of building systems. The role requires expertise in various...Hourly payMinimum wageApprenticeshipWork at officeLocal areaFlexible hoursShift workWeekend workAfternoon shift$120.5k - $231k
...communities and building trust in how we show up, everywhere & always. Want in? Join the #VTeamLife. What You'll Be Doing As the Principal Engineer for Adobe Experience Platform (AEP) you are the technical architect, strategy and operational anchor of the Verizon Business Group...Full timeTemporary workPart timeWork experience placementWork at officeImmediate startRemote workWork from homeShift work3 days per week- ...-oriented, creative problem solver, and collaborative Solutions Engineer to join our Integrated Solutions team. The Solutions Engineer plays... ..., conduct hospital walk‑throughs, co‑lead in person working on site sessions. Specific vision abilities required by this job include...Work experience placement
$105k - $160k
...Job Purpose and Impact The Senior Professional, Platform Engineering job designs, develops and maintains digital technology infrastructure... ...to automate the deployment process, ensuring smooth and reliable releases. COLLABORATION: Partners with cross functional...Work experience placement$128k - $216k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe your... ...Fiserv. Job Title Senior AI Solutions Engineer About the role As a Senior AI... ...shaping the future of fintech, this role is on-site Monday through Friday This role...Work at officeVisa sponsorshipMonday to Friday- ...ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Job Overview As an Solution Engineer, you will be a key player in shaping the success of our product by collaborating with the sales and product development teams. You...Work experience placement
- ...Job Title: Platform Engineer / Solution Architect Location: Alpharetta, GA (Hybrid - 3 Days Onsite) Type: Contract Required Technical Skills: Strong experience architecting and deploying Kubernetes-based microservices and containerized applications...Contract work
- ...Job Title: OpenShift Platform Engineer Location: Alpharetta, GA (Hybrid - Partially Onsite) Job Description: We are seeking an experienced OpenShift Platform Engineer to support, manage, and enhance enterprise-grade Red Hat OpenShift environments...
$120k - $160k
...We reach higher. We do the right thing—today and for generations to come. Job Purpose and Impact ~ The AI Security Engineering Manager will help solidify foundation for the company's modern business applications. In this role, you will apply your knowledge...- ...value. Job Summary Greenstone is seeking a mid-level Software Engineer to join our integrations team. This role is ideal for someone... ...– 30% Monitor, troubleshoot, and optimize the performance and reliability of integration flows. Ensure integration meets security, auditability...Remote work
- ...Developer to join our ever-evolving Heartland School Solutions Engineering Team and help shape the future of global commerce. In this... ...K-12 schools, helping clients streamline operations, improve reliability and deliver better outcomes for the communities they serve. Your...
- ...primarily .NET and SQL Server), learn our delivery processes (Azure DevOps, CI/CD), and collaborate with senior engineers and business stakeholders to ship reliable, secure features. This role is a strong fit for someone who has solid fundamentals and wants to grow quickly...Work experience placementH1bWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- IT site lead Alpharetta, GA
- junior website developer Alpharetta, GA
- site safety Alpharetta, GA
- site services specialist Alpharetta, GA
- site leader Alpharetta, GA
- on-site clinical research associate (traveling/remote) Alpharetta, GA
- site reliability engineer
- junior site reliability engineer
- site reliability engineer remote
- site reliability engineering manager



