Lead Site Reliability Engineer
Luxoft
Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.
Responsibilities
Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
Identify reliability, operational, and technology risks requiring escalation to management.
Promote an environment that supports a culture of belonging and reflects the Client brand.
Maintain Client internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
Complete other related duties as assigned.
Skills
Must have
Strong experience in observability and monitoring, including hands-on expertise with:
Dynatrace
OpenTelemetry (OTel)
Distributed tracing
Metrics collection and analysis
Centralized logging and log aggregation
Alerting and dashboard development
Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
Strong proficiency in Infrastructure as Code (IaC) using Terraform.
Experience with CI/CD pipelines, deployment automation, and operational tooling.
Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
Cloud & Platform Expertise
Strong experience with Microsoft Azure, including:
Azure App Services
Resource Groups
Azure networking concepts
Scaling and performance optimization
Deployment and release management
Application lifecycle management
Experience leveraging Azure-native operational tooling such as:
Azure Monitor
Application Insights
Log Analytics
Azure dashboards and alerting
Experience supporting cloud-native and hybrid infrastructure environments.
Reliability & Engineering Practices
Demonstrated experience implementing and operating SRE practices, including:
Service Level Objectives (SLOs)
Service Level Indicators (SLIs)
Error budgets
Incident management
Problem management
Root Cause Analysis (RCA)
Reliability automation
Ability to improve system reliability through:
Performance tuning
Capacity planning
Observability-driven insights
Proactive issue detection
Reliability engineering initiatives
Experience developing automated recovery mechanisms and self-healing solutions.
Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
Nice to have
Experience supporting large-scale enterprise applications in regulated environments.
Strong analytical and troubleshooting skills related to production systems and distributed architectures.
Experience working in Agile and DevOps operating models.
Ability to work autonomously and lead complex reliability initiatives.
Strong organizational and time management skills.
Advanced verbal and written communication skills.
Experience driving project milestones and delivery commitments.
Proven experience leading major incident response and post-incident improvement efforts.
Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.
Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferred.
Other
Languages
English: C1 Advanced
Seniority
Lead
$139.7k - $232.9k
...implementing, and continuously improving highly reliable, scalable, and resilient platform... ...as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering... ...the Software Development Lifecycle. Leads complex initiatives, influences enterprise...SuggestedFull timeWork experience placement- Job Title: Site Reliability Engineer AWS Location: NY Work Mode: Onsite/Hybrid Engagement: C2C Contract Experience: 6+ Years Job Summary We are seeking an experienced Site Reliability Engineer with strong Amazon Web Services expertise to build...SuggestedContract work
$92.7k - $125.2k
This OpportunityWSP is currently seeking a Lead Professional, Biology to join the team in New York at either our Buffalo, Syracuse, or... ...information on a particular environmental problem or issue, verify site characteristics, and/or to plan for future environmental needs....SuggestedWork at officeLocal area$95k - $175k
...To:Manager, Contract AdministrationWork Schedule:Hybrid - Santa Barbara, CAOverallMoog Military Aircraft LLC. is seeking a Contracts Lead for our Turreted Weapon Systems Business Unit, Defense Division.Moog requires an experienced and dynamic Contracts Lead to advance...SuggestedFull timeContract work$110k - $125k
...Local - Buffalo, NYMoog’s Corporate HR Technology Team is looking for an experienced, detail-oriented Workday Benefits and Compensation Lead to join them!As the Workday Benefits and Compensation Lead, you will serve as the primary functional expert for Moog’s global...SuggestedFull timeWork at officeLocal areaRemote workFlexible hours$89.6k - $149.3k
...education and/or work experience, including a minimum of 5 years’ relevant work experienceStrong attention to detailStrong ability to lead teamsExcellent project management skillsStrong analytical abilityCreative problem-solving skillsExcellent interpersonal skills...Full timeWork experience placement$116.4k - $194k
...certification strategies.Ensure governance models are tool aware but tool agnostic, scalable, and not overly reliant on manual effort.Lead governance strategy for cloud identity, including Azure AD / Entra ID, RBAC, conditional access, and privileged role governance....Full timeWork experience placementWork from home- ...A leading software firm is seeking a Talent Acquisition Partner II to enhance their payroll operations and support HR compliance. The successful candidate will lead payroll processes, ensure accuracy, and maintain state and local tax compliance. Candidates should have...Full timeLocal area
$70k - $85k
...can mean deeper job satisfaction, better rewards, and a great quality of life inside and outside of work. Job Title : Site Repair Lead Reporting To: Manager, Global Repair Work Schedule: Onsite - Buffalo, NY Our team in Commercial Aircraft Group...Relocation packageFlexible hoursShift workDay shift$95.92k - $222.03k
...society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched... ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Temporary workWork at officeRemote workFlexible hours$95.92k - $222.03k
...society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched... ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Temporary workWork at officeRemote workFlexible hours- ...delivering a true turn-key product to our customers. Our scope spans Engineering, prototype builds, through full production, requiring teams... ...Customer Obsession I Speed I Agility Overview The Working Lead for our 2nd Shift Welding team is responsible for performing...Contract workLocal areaMonday to FridayShift workAfternoon shift
$95.92k - $222.03k
...specialist skills to ensure that all the engineering services are operated, maintained and... ...responsible innovation. We are one of the world's leading AI and digital infrastructure providers,... ...locally to NTT DATA offices or client sites. This ensures we can provide timely and...Temporary workWork at officeRemote workFlexible hours- Delaware North Sportservice is seeking a seasonal Alcohol Compliance Supervisor to join our team at KeyBank Center in Buffalo, New York. The role focuses on ensuring compliance with state, federal, and local alcohol laws as well as company procedures. You will supervise...Seasonal workLocal areaNight shiftWeekend work
- ...become part of a global team of over 50,000 planners, designers, engineers, scientists, digital innovators, program and construction... ...for cornerstone municipal utilities in Western New York, while leading the local practice across wastewater and water treatment, conveyance...Work at officeLocal areaWorldwideRelocation packageFlexible hours
$139.7k - $232.9k
...day per week.Overview: The Senior Principal Cloud & Platform Engineer is a senior‑level individual contributor responsible for defining... ...platform standards are aligned with enterprise architecture, reliability, and risk management expectations.Communicate technical...Full timeWork from home1 day per week$80.9k - $134.8k
...Configures, supports, and enhances Microsoft 365 collaboration and productivity solutions to ensure operational continuity, platform reliability, and a positive user experience. Serves as a primary resource for Microsoft 365 service requests, consultations, intake...Full timeWork experience placementFlexible hours$25 - $50 per hour
...Role Overview TSA is accepting applications for Lead and Supervisory Transportation Security Officers at airports in Buffalo. These roles are ideal for individuals looking to step into leadership positions within airport security operations. TSA provides training to...Shift workNight shiftWeekend work$160k - $180k
...Lead Architect Location: Buffalo, NY, United States Salary Range: USD 160000-180000 Annually Date Posted:May 19, 2026 Description... .... What you bring: Bachelor's degree in Computer Science, Engineering, Business or related discipline; equivalent years of experience...Contract workRemote work- ...Vas Lead As a Vas Lead, one will be responsible for unpacking and repacking products directly from customers to assemble various displays and differing case sizes per customer specification, picking orders, and helping to lead the staff within the department. Duties...Monday to Friday
$15.11 - $18.89 per hour
...Job Description Zone Lead FT Pay Range: $15.11 - $18.89 The Zone Lead (ZL) reports to the Assistant Store Manager and supports... ...to detail, e.g., standards, processes, marketing, etc. Reliable and trustworthy role model for ethical and honest behavior. Does...Flexible hoursShift workNight shift- .... We are committed to being America's best first job. Let's talk. Make your move. See a day in the life of a Guest Experience Lead at McDonald's Requirements: We believe in letting you do you. If you're looking for a part-time job that supports your full-time...Full timePart time
$150.16k
...Job Description CDM Smith is seeking a Lead Enterprise Architect to join our Corporate Business Technology team. In this role, you will be responsible for designing and governing the firm’s enterprise application architecture, ensuring alignment with our strategic IT...H1b- ...:ACV Auctions is looking for a Software Engineer to join our MAX Engineering team and work... ...maintenance, with a strong focus on reliability, performance, and operational excellence... ...; data, enterprise applications, and IT.Lead technical design and architecture decisions...
- What you will do:The Software Engineer III will work on a cross-functional agile team in operational tools, product, or infrastructure... ...scalable services, plan feature roll-out, and ensure high reliability and performance for your products.Conduct code reviews, contribute...
$120k - $150k
...exceptional things. A business that employs over ten thousand engineers and technologists with a small company feel. An enterprise where... ...architecture and interface definition, and requirements validation. Lead the design process and provide technical leadership within Moog...Full timeRelocation packageFlexible hours$76.1k - $126.9k
...in C++ and data store which serves as the cash flow projection engine underlying Yield Book’s analytics for mortgage-backed and asset... ...make a significant impact, we'd love to hear from you.LSEG is a leading global financial markets infrastructure and data provider. Our...Full timePart timeInternship$97.1k - $161.8k
...the expert level for mainframe software engineering, technical solution delivery, and technical... ...as an individual-contributor technical lead for critical mainframe applications and... ...monitoring, support documentation, and reliability improvements. Apply AI-assisted engineering...Full timeWork experience placement$70 - $75 per hour
Software Engineer II (Java/Azure) Location: Buffalo, NY (Hybrid - 4 Days Onsite)Employment Type: W2 Only*OPEN TO RELOCATION CANDIDATES*... ...by law.About TEKsystems and TEKsystems Global Services We’re a leading provider of business and technology services. We accelerate business...Contract workTemporary workWork experience placementRelocation$16.5 - $17.5 per hour
Title: Custodial Lead Job Description: Job Overview The Custodial Lead will be responsible for the cleanliness and sanitation of the areas assigned and provides some work direction to custodial staff. Roles & Responsibilities To perform this job successfully and safely...Hourly payImmediate startMonday to FridayShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Buffalo, NY
- lead operating engineer Buffalo, NY
- on-site clinical research associate (traveling/remote) Buffalo, NY
- website coordinator Buffalo, NY
- junior website developer Buffalo, NY
- site leader Buffalo, NY
- historic site Buffalo, NY
- construction site safety Buffalo, NY
- official site Buffalo, NY
- site services specialist Buffalo, NY



