GenAI Platform Engineering Lead
$116.4k - $194kM&T Bank
Manages the activities of Engineering Team Leaders, Engineering Supervisors and/or engineering units responsible for the reliability, observability and operational readiness of the Bank’s AI platform. Provides day-to-day direction for the teams and applications in alignment with departmental goals and the needs of the clients they support.Serves as the technical lead for AI operations and service assurance, with accountability for the operational layer of the AI platform. Works closely with the Platform Engineering Manager, who owns the broader platform, to ensure production systems are secure, resilient, observable, cost-effective and operationally ready. Oversees incident response practices, runbooks, monitoring, infrastructure automation and production support.Responsible for managing client relationships and expectations, prioritizing the project queue and achieving individual and organizational objectives at minimum cost.Primary ResponsibilitiesLead the reliability, observability and operational readiness of the AI platform, including production monitoring, incident response, service assurance, infrastructure automation and operational controls.Establish and maintain operational runbooks, escalation procedures, service-level indicators, service-level objectives and incident response practices.Measure operational performance through platform availability, incident response time, mean time to recovery, model latency, token consumption and infrastructure cost.Provide leadership during production incidents. Coordinate troubleshooting, communication, escalation, root-cause analysis and corrective actions.Build and maintain observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure Monitor, Log Analytics and/or comparable technologies.Oversee the deployment and operation of Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID and infrastructure managed through Terraform.Promote infrastructure-as-code and automated deployment practices using Terraform, GitHub Actions, GitLab CI and/or comparable tools. Ensure platform changes are deployed through controlled pipelines rather than manual processes.Drive automation of recurring operational activities using Python, Bash, PowerShell and other appropriate scripting technologies.Oversee AI-specific operational capabilities, which may include token cost tracking, model latency monitoring, provider failover, caller-level rate limiting, prompt logging, personally identifiable information interception and infrastructure-layer content filtering.Partner with cybersecurity and risk teams to support Security Information and Event Management integrations and security event feeds from application infrastructure.Manage and participate in consultations with client management to analyze short-range business requirements and recommend innovations that anticipate the future impact of changing business and technology needs. Build and maintain positive client relationships.Monitor technology direction, industry trends and vendor applications related to AI platforms, site reliability engineering, cloud infrastructure, observability and service assurance.Research and initiate changes to existing processes, technologies and operating models when necessary.Lead vendor and product analysis and provide recommendations.Oversee application development support, testing efforts, technology infrastructure, project management and other assigned technology domains.Serve as a subject matter expert for AI platform operations, production reliability, observability and service assurance.Build rapport across the organization and maintain a professional level of communication and cooperation with technology, business, risk, cybersecurity and vendor partners.Maintain relationships with vendors and professional organizations.Direct team activities, assign personnel to projects and provide technical and operational guidance.Ensure schedules and commitments are completed. Lead short-term staffing and capacity planning.Implement technology consistent with Division standards and long-range plans. Ensure adherence to Department and Technology standards and procedures, including documentation, audit trail and change management requirements.Translate business and operational requirements as needed to assist staff in preparing detailed specifications for system enhancements.Evaluate and manage recommended designs based on business, operational and technology requirements. Identify, communicate and resolve issues and concerns.Manage project plans and coordinate major project and production-readiness activities. Remain current on work outside the team that may affect the team, platform or client environment.Develop and manage multiple cost center budgets, including cloud infrastructure and AI platform operating costs.Recommend and implement policies and procedures that improve the performance, reliability and effectiveness of the Department.Exercise the usual authority of a manager concerning staffing, performance appraisals, promotions, salary recommendations, performance management and terminations.Understand and adhere to the Company’s risk and regulatory standards, policies and controls in accordance with the Company’s Risk Appetite. Design, implement, maintain and enhance internal controls to mitigate risk on an ongoing basis. Identify risk-related issues requiring escalation to management.Promote an environment that supports belonging and reflects the M&T Bank brand.Maintain M&T internal control standards, including the timely implementation of internal and external audit points and resolution of issues raised by external regulators, as applicable.Complete other related duties as assigned.Scope of ResponsibilitiesOversees a team where the majority of employees are engineers, architect individual contributors, Engineering Supervisors and/or Engineering Team Leaders. Leads the operational capabilities supporting the AI platform and partners closely with the Platform Engineering Manager, application teams, cybersecurity, risk, architecture and other technology stakeholders.This role requires both people leadership and technical depth. The manager is expected to provide hands-on technical direction, support complex production troubleshooting and ensure the team’s operational practices meet the Bank’s reliability, security, risk and regulatory expectations.Supervisory/Managerial Responsibilities5 to 10Skills and Success IndicatorsOperational ownership: Measures success through platform availability, incident response time, mean time to recovery, operational readiness and the quality of documented runbooks.Technical depth: Can review telemetry, interpret traces, troubleshoot infrastructure and gateway configurations and provide technical direction during complex production issues.Leadership: Develops engineers and leaders, establishes clear accountability and creates an environment that supports collaboration, belonging and continuous improvement.Automation mindset: Identifies repeatable operational activities and moves them toward scripted, pipeline-based and self-service solutions.Cost awareness: Treats token consumption, infrastructure usage and cloud expense as core operational metrics.Risk and control orientation: Builds auditability, security, change management and regulatory expectations into operational processes.Client partnership: Builds trust with business and technology stakeholders through clear communication, dependable execution and transparent incident management.Education and Experience RequiredA combined minimum of 9 years’ higher education and/or work experience, including a minimum of 4 years’ engineering and/or architecture experience and 3 years leadership experienceMinimum of 5 years’ experience in site reliability engineering, platform operations, infrastructure engineering or a related discipline, including experience supporting production systems through an on-call modelHands-on experience building, operating and troubleshooting Azure infrastructure and services, including Azure API Management, Azure Monitor, Log Analytics, Microsoft Entra ID and TerraformExperience building and supporting observability pipelines using OpenTelemetry, Prometheus, Grafana, Azure-native technologies and/or comparable toolsStrong understanding of metrics, traces and logs, including when and how each should be used to monitor and troubleshoot production systemsExperience implementing continuous integration and continuous delivery practices for infrastructure using Terraform, GitHub Actions, GitLab CI, infrastructure-as-code patterns and/or comparable technologiesStrong scripting and automation skills using Python, Bash, PowerShell and/or comparable languagesExperience leading or supporting production incident response, including troubleshooting, escalation, root-cause analysis and service restorationCapable of working on multiple projects of a complex natureProficiency with project management, word processing and spreadsheet applicationsComplete understanding of the system development life cycleExcellent problem-solving skills to assist in issue resolutionFamiliarity with application development software, cloud infrastructure and hardware platformsExcellent verbal and written communication skillsExcellent analytical and decision-making skillsStrong project management and presentation skillsExperience encouraging teamwork and serving as a role model when leading and directing othersUnderstanding of the technical, business, operational, risk and cost impacts of a project, platform or production issueEducation and Experience PreferredBachelor’s degreeMinimum of 10 years’ technology management, site reliability engineering, platform operations or large program leadership experienceExperience operating AI or large language model platforms in a production environmentExperience with large language model operational concerns, including token cost tracking, model latency monitoring, provider failover and rate limiting by caller identityExperience integrating application infrastructure with Security Information and Event Management platforms and security event feedsFamiliarity with AI gateway patterns, including prompt logging, personally identifiable information interception and infrastructure-layer content filteringExperience working in a regulated environment where audit trails, access controls and change management practices are requiredExtensive application and product knowledge within AI platforms, cloud infrastructure, site reliability engineering, observability and/or service assuranceSubject matter expert understanding of supported applications, with advanced knowledge of interfacing and integrated applicationsUnderstanding of multiple business areas and their functionsProven mentoring and leadership capabilitiesExperience with the technologies, applications and functions of the area being ledGood understanding of the Bank’s application frameworkAwareness of the Bank’s business plan and strategic objectives, with the ability to help shape directionSelf-motivated with the ability to motivate and develop othersUnderstanding of the supported businesses and their terminologySkills and Success IndicatorsOperational ownership: Measures success through platform availability, incident response time, mean time to recovery, operational readiness and the quality of documented runbooks.Technical depth: Can review telemetry, interpret traces, troubleshoot infrastructure and gateway configurations and provide technical direction during complex production issues.Leadership: Develops engineers and leaders, establishes clear accountability and creates an environment that supports collaboration, belonging and continuous improvement.Automation mindset: Identifies repeatable operational activities and moves them toward scripted, pipeline-based and self-service solutions.Cost awareness: Treats token consumption, infrastructure usage and cloud expense as core operational metrics.Risk and control orientation: Builds auditability, security, change management and regulatory expectations into operational processes.Client partnership: Builds trust with business and technology stakeholders through clear communication, dependable execution and transparent incident management.#LI-JB3M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $116,400.00 - $194,000.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.LocationBuffalo, New York, United States of AmericaSummaryLocation: Buffalo, NYType: Full time
$139.7k - $232.9k
OverviewManages the activities of several Engineering and/or Architecture Managers and/or Team... ...and operating M&T Bank’s enterprise AI platform. This team owns the Azure API Management... ...internal teams and hires, develops and leads the engineering talent responsible for the...SuggestedFull timeTemporary workWork experience placement$116.4k - $194k
OverviewResponsible at the expert level for transforming engineering, operational, security, and quality data... ...and software delivery outcomes. Leads the collection, integration, analysis,... ...Lifecycle (SDLC), observability platforms, incident management systems, production...SuggestedFull timeWork experience placement$139.7k - $232.9k
...Seneca One Buffalo, NY location, with the flexibility to work from home one day per week.Overview: The Senior Principal Cloud & Platform Engineer is a senior‑level individual contributor responsible for defining, standardizing, and evolving core cloud and platform...SuggestedFull timeWork from home1 day per week$63k - $140k
...Associate Job Description & Summary The Opportunity As a GenAI Python Systems Engineer – Experienced Associate, you will leverage advanced... ...aligned to data engineering, machine learning, and cloud platforms, including AWS, Google Cloud, Microsoft Azure, Databricks...SuggestedFull timeH1b$77k - $202k
...Associate Job Description & Summary The Opportunity As a GenAI Python Systems Engineer – Senior Associate, you will play a pivotal role in... ...aligned to data engineering, machine learning, and cloud platforms, including AWS, Google Cloud, Microsoft Azure, Databricks...SuggestedFull timeH1b$116.4k - $194k
Job SummarySDLC GenAI Automation & Tooling Integrations Engineer will play a key role in automating and modernizing the... ...with the SDLC Program Governance Lead, SDLC BSAs, GenAI engineering... ...Power BI, and related engineering platforms.Automate SDLC artifact generation...Full timeWork experience placementWork at office$116.4k - $194k
Role SummaryThe Lead AI Platform Engineer is a hands‑on technical contributor responsible for designing, developing, and integrating Generative... ...combines strong software engineering fundamentals with applied GenAI experience and partners closely with product, architecture,...Full timeWork experience placement- ...AI & Platform Software EngineerLocation: Remote within the USA, or onsite... ...Canada.Project Description A leading financial services client is building its enterprise GenAI capability across two closely... ...and tooling, quality engineering and test data management, and...Contract workFor contractorsWork at officeRemote work
$98.3k - $175.23k
WSP is seeking a Renewable Energy Electrical Engineering Lead for our Buffalo, NY office. The following locations will also be considered: New York NY, Chicago IL, Boulder, CO and Austin, TX. This opportunity involves being part of our vastly experienced and knowledgeable...For subcontractorWork at officeLocal area- ...become part of a global team of over 50,000 planners, designers, engineers, scientists, digital innovators, program and construction... ...talented Structural Dam Safety Engineering Project Manager/Technical Lead to join the East Region Dams Team. This position will support a...Part timeFor subcontractorSecond jobWork at officeLocal areaWorldwideRelocationFlexible hours
- Job DescriptionGestalte die digitale Transformation der Öffentlichen Verwaltung!Konzeption, Entwicklung und Implementierung eines Lösungsbausteins für Network Functions Virtualization (NFV) auf Basis von Open-Source-TechnologienAblösung traditioneller, hardwarebasierter...Home office
$99k - $232k
...Transformation journey, utilizing Oracle Cloud ERP and EPM to create the next generation Finance function. As a Manager, you will lead teams and manage client accounts, focusing on strategic planning and mentoring junior staff. You are accountable for project success...H1b$116.4k - $194k
The Lead Engineer - Enterprise Metrics Platform, Data Engineering & Analytics will architect, build, and operate the enterprise technology metrics platform responsible for aggregating, transforming, governing, and visualizing data from multiple technology systems into a...Full timeWork experience placement$168k - $230k
...AI to life and is the company behind Astro, the industry‑leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable... ...‑source software. We’re looking for a Senior Software Engineer to join our Platform Engineering team. You get to go in...$128.9k - $214.9k
...initiatives at the division and department level. Leads design, planning and implementation for... ...Technology, Finance, Architecture, Engineering, and Business teams to drive cost... ...optimization, and value realization across cloud platforms and AI services. • Develop and maintain...Full timeWork experience placement$110k - $155k
...satisfaction, better rewards, and a great quality of life inside and outside of work.Job Title:Senior Cloud Engineer - Platform Automation & DevOpsReporting To:Lead Analyst., Project/SystemWork Schedule:Hybrid - Buffalo, NYMoog's Corporate Group is looking for a Senior...Full timeLocal areaRelocationFlexible hours3 days per week$93.58k - $142.64k
...DevOps EngineerCUBRC is seeking an experienced DevOps engineer to join our team located in Buffalo, NY. As a DevOps engineer, you will... ...are expected to have a working knowledge of modern DevOps platforms, tools, and processes including their application to both locally...Work at office- DevOps EngineerShould have good experience in DevOpsImplement and support Continuous Integration and Deployment Pipelines. ImplementationsExperience in Agile development practicesExperience in GIT version controllingShould have experience in enhancing your team member’...
- Cloud Competency Center, Sr. Azure Cloud ArchitectWork Location: Buffalo/Cheektowaga, NY (100% Remote) Ability to Work Remote: Yes, candidate should be located in Eastern Time Zone, US (client will also consider someone in Central Time Zone open to working in Eastern Time...Immediate startRemote work
$116.4k - $194k
This role sits on a high‑visibility, agile engineering team building a next‑generation digital banking platform from the ground up. You’ll join a tight‑knit team of 5-... ...over design, delivery, and technical direction.As a Lead Software Engineer, you’ll work hands‑on across...Full timeWork experience placement$125.5k - $261.6k
...help to build a better working world. The opportunity We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY’s AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI...Full timeContract workSummer holidayFlexible hours- ...We are seeking a skilled and motivated Mid-Level DevOps Engineer with expertise in Linux and Windows environments, automation,... ..., pipelines, and applications. • Use monitoring and logging platforms to analyze system health and performance, troubleshoot issues,...
$139.7k - $232.9k
...highly reliable, scalable, and resilient platform solutions across the enterprise.... ...matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices... ...across the Software Development Lifecycle. Leads complex initiatives, influences enterprise...Full timeWork experience placement$124k - $280k
...Oracle Cloud ERP and EPM, along with emerging technologies like RPA, Machine Learning, and Analytics. As a Senior Manager, you will lead large projects and innovate processes, focusing on achieving results and maintaining operational excellence. You will interact with...H1b$125.5k - $261.6k
...build a better working world. The opportunity We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. These are the systems that ship, run, and make fully visible every AI workload....Full timeSummer holidayFlexible hours- ...Position Title: Site Reliability Engineer Location: Buffalo, NY Duration: 6 Months with possible extension Lead Site Reliability Engineer (Level 61 IC) Overview... ...excellence of critical banking platforms and applications. Serves as a senior individual...Work experience placement
- ...About the job Lead Site Reliability Engineer Job Title: Lead Site Reliability Engineer Location: Remote within the USA, or onsite... ...performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will...Work at officeRemote work
$142.6k - $297.2k
...implementations along with providing experience in leading practices, methods, and resources in the... ..., Dell Boomi or any leading integration platform to architect, design and build... ...in Computer Science, IT, Computer Engineering, MIS, Mathematics, or related field (MBA...Full timeSummer holidayFlexible hours$146.5k - $219.7k
...for the functional management of approximately ~8-10 software engineering personnel assigned to a variety of programs. Duties include conducting... ....Demonstrated coaching skills are a plus.Previous experience leading a team of 3-5+ employees with a record of on-time performance....Full timeRelocation packageShift work$121k - $181.4k
...for the functional management of approximately ~8-10 software engineering personnel assigned to a variety of programs. Duties include conducting... ....Demonstrated coaching skills are a plus.Previous experience leading a team of 3-5+ employees with a record of on-time performance....Full timeRelocation packageShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GenAI Platform Engineering Lead. Be the first to apply!

