Lead ML Platform Engineer (SRE / FTE / Onsite)
$83.52k - $125.28kNTT Data
Req ID: 388174
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to join our team in Charlotte , North Carolina (US-NC) , United States (US) .
Job Duties and Responsibilities:
The Lead ML Platform Engineer provides architecture and hands-on engineering leadership for the Cortex Predictive AI Platform across cloud and on-premises environments. This role will establish and implement reusable, secure, scalable standards that enable data scientists, ML engineers, and application teams to build, validate, deploy, monitor, and operate predictive models efficiently and reliably.
The successful candidate will lead technical design and engineering decisions across the ML platform lifecycle, including governed data and feature access, model development environments, training and validation workflows, model delivery pipelines, real-time and batch inference, observability, reliability, and operational readiness. This role will also mentor engineering teams and transfer knowledge to support sustainable platform operations and adoption.
Key Responsibilities
- Define and lead the target architecture for predictive AI and ML platform capabilities spanning public cloud and on-premises environments.
- Design, build, and operate reusable platform services supporting the end-to-end ML lifecycle: governed data and features, model development, training, validation, deployment, inference, monitoring, and operations.
- Establish scalable reference architectures, engineering standards, reusable templates, and implementation patterns for ML workloads across the Cortex portfolio.
- Lead platform engineering for GCP and multi-cloud environments, including secure connectivity, identity, network controls, compute, storage, and managed AI/ML services where applicable.
- Design and operate Kubernetes-based ML platforms using GKE, OpenShift, and associated container, workload orchestration, and resource-management capabilities.
- Implement and improve MLOps capabilities for experiment tracking, model packaging, validation, approval gates, model registry integration, deployment automation, rollback, and lifecycle management.
- Build CI/CD pipelines and infrastructure automation for platform services, ML workflows, model delivery, and environment provisioning.
- Enable model migration from legacy environments into standardized Cortex platform patterns, minimizing delivery risk and operational disruption.
- Engineer production-grade real-time and batch inference capabilities, including API-based serving, scalable runtime patterns, resiliency, performance, and operational support.
- Partner with data engineering, data governance, security, privacy, risk, model validation, and application teams to ensure data protection and control requirements are embedded into platform design.
- Implement platform observability, including logs, metrics, traces, dashboards, alerts, service-level indicators, service-level objectives, and operational runbooks.
- Drive reliability engineering practices for ML platform services, including capacity planning, high availability, disaster recovery, incident management, root-cause analysis, and continuous improvement.
- Ensure platform designs meet enterprise security requirements for authentication, authorization, secrets management, encryption, data access, auditability, and environment isolation.
- Provide technical leadership, architecture reviews, code reviews, design guidance, and mentoring to ML platform engineers and adjacent delivery teams.
- Produce clear technical documentation, reference implementations, operational procedures, and knowledge-transfer materials to enable self-service adoption and long-term support.
Required Qualifications
- 8+ years of experience in platform engineering, cloud engineering, infrastructure engineering, SRE, MLOps, or related technical roles.
- 4+ years of experience designing, building, or operating enterprise AI/ML or data platforms.
- Demonstrated experience leading architecture and engineering delivery for complex, production-grade cloud and/or on-premises platforms.
- Strong hands-on experience with GCP and working knowledge of multi-cloud or hybrid-cloud architecture.
- Experience with Kubernetes-based platforms, including GKE and OpenShift, in production environments.
- Strong experience implementing MLOps capabilities, model lifecycle workflows, or ML platform services.
- Proficiency in Python for platform automation, integration, operational tooling, or ML workflow development.
- Experience with CI/CD, Git-based development, automated testing, deployment automation, and infrastructure-as-code practices.
- Strong understanding of enterprise security, data protection, identity and access management, secrets management, encryption, audit logging, and secure software delivery.
- Experience implementing observability, monitoring, alerting, dashboards, SLOs, incident response, and operational runbooks.
- Experience mentoring engineers and communicating technical architecture decisions to engineering, product, security, data, and executive stakeholders.
Required Skills / Knowledge
- Enterprise ML platform architecture and end-to-end predictive model lifecycle management.
- GCP, hybrid cloud, multi-cloud, on-premises platform, networking, identity, and security concepts.
- Kubernetes, GKE, OpenShift, containers, workload orchestration, and scalable compute platforms.
- MLOps, model development environments, model registries, validation workflows, model deployment, and model monitoring.
- Python, CI/CD, Git, automated testing, infrastructure automation, and API-based integration.
- Real-time and batch inference architecture, model-serving patterns, performance optimization, and operational support.
- Data protection, governance, access controls, encryption, auditability, and regulated-platform design.
- Observability, telemetry, dashboards, alerting, SLI/SLO design, reliability engineering, and production troubleshooting.
- Technical leadership, reusable pattern development, engineering documentation, and knowledge transfer.
Preferred Qualifications
- Experience with Vertex AI or comparable cloud ML platform services.
- Experience designing or operating on-premises AI/ML platforms, private cloud, or hybrid ML workloads.
- Experience with feature stores, model registries, experiment tracking, data lineage, model governance, or model risk-management processes.
- Experience supporting model migration, platform modernization, or transition from legacy data science and ML environments.
- Experience with real-time, low-latency model-serving systems and event-driven inference architectures.
- Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation and deployment tooling.
- Experience in banking, financial services, healthcare, insurance, or another regulated enterprise environment.
- Experience establishing self-service platform capabilities for data scientists, ML engineers, and application teams.
Expected Outcomes
- A secure, scalable, and reusable Cortex ML platform architecture spanning public cloud and on-premises environments.
- Standardized MLOps, CI/CD, and model-delivery patterns that reduce time to train, validate, deploy, and operate predictive models.
- Reliable platform capabilities for governed data and features, model migration, batch and real-time inference, and production operations.
- Improved observability, resiliency, service-level management, and operational readiness for ML platform services and models.
- Reusable engineering standards, reference implementations, documentation, and knowledge-transfer assets that enable self-service adoption and sustainable platform support.
#LI-NorthAmerica
NTT DATA provides a reasonable range of compensation for U.S.-based positions. The starting pay range for this role is $83,520.00 - $125,280.00. Actual compensation will depend on a number of factors, including the candidate’s relevant experience, technical skills, and other qualifications.
This position may also be eligible for incentive compensation based on individual and/or company performance.
This position is eligible for company benefits including medical, dental, and vision insurance with an employer contribution, flexible spending or health savings account, life and AD&D insurance, short and long term disability coverage, paid time off, employee assistance, participation in a 401k program with company match, and additional voluntary or legally-required benefits.
About NTT DATA
NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries. We also offer clients access to a robust ecosystem of innovation centers as well as established and start-up partners. NTT DATA is a part of NTT Group, which invests over $3 billion each year in R&D.
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com, @nttdatafed.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us form,
NTT DATA endeavors to make accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at . This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here . If you'd like more information on your EEO rights under the law, please click here . For Pay Transparency information, please click here .
- ...operating cloud-native applications and platforms on AWS and Azure, including VPC/VNet, IAM... ..., model serving, embeddings, prompt engineering, AI evaluations, and agentic frameworks.... ...management, observability, monitoring, logging, SRE practices, and production support....SuggestedTemporary work
- ...Reference Number: 25-00760Title: AWS Python ML Developer - Lead LevelPosted Date: 2025-07-10Company:... ...: Malvern, PA (Hybrid- 3 days a week onsite & must be willing to relocate and be onsite... ...build of an internal virtual advisor platform (similar to a ChatGPT-style copilot for...SuggestedFull timeRelocation3 days per week
- ...Title : GCP DATA Platform Engineer Location: Charlotte, NC Onsite Job Type: Contract About The Role We are building a hybrid data platform that bridges an On-premises big data cluster with Google Cloud Platform and we need an engineer who can own...SuggestedContract work
- ...enterprise-grade MLOps/GenAI platforms across AWS and Azure, leveraging... ...and automated cloud-native ML infrastructure using Terraform... ...implementing model evaluation, prompt engineering, state management, caching,... ..., logging, monitoring, SRE practices, performance tuning,...SuggestedContract work
- Charlotte, North CarolinaHybridFull Time$120k - $160kSenior AI/ML Engineer (Python)Onsite — Wilmington, NC Our client, a fintech company building... ..., or a quantitative discipline Familiarity with cloud platforms and model-serving/MLOps infrastructure Benefits • Medical...SuggestedFull time
- ...We are seeking an experienced Lead Alteryx Administrator / BI Platform Engineer with 10+ years of experience administering, deploying, and supporting enterprise Business Intelligence and Analytics platforms. The ideal candidate will have deep expertise in Alteryx Server...Temporary work
- ...Charlotte, NC-based opportunity is for a full-time Associate Platform Engineer working onsite with a fast-growing organization building enterprise-grade... ...as code (CloudFormation or Terraform) Exposure to AI/ML systems, LLMs, or agent-based architectures Experience building...Full time
- A social impact organization based in Charlotte, NC is seeking a Freelance Event Specialist to manage logistics for in-person events. The ideal candidate possesses strong communication skills and has experience in both large-scale and intimate events. Responsibilities ...Hourly payFreelance
$87.12k - $181.5k
...Python Backend Developer - Hybrid/Onsite to join our team in Charlotte... ...code following software engineering best practices.Perform unit testing... ...with cloud platforms such as AWS, Microsoft Azure,... ...mentoring junior developers or leading small technical initiatives.Python...Full timeTemporary workWork at officeRemote workFlexible hours$106.84k - $160.25k
...now.We are currently seeking a Data Engineer - Cybersecurity Analytics - Onsite to join our team in Charlotte,... ...high-performing data pipelines and platforms.Key ResponsibilitiesDesign, develop... ...innovation. We are one of the world’s leading AI and digital infrastructure...Full timeTemporary workWork at officeRemote workFlexible hours- ...Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the... ...enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical... ...infrastructure, cloud, platform engineering, or SRE experience. ~5+ years managing...Work at officeFlexible hoursShift workDay shift
- ...Meghana GorusuCompany: SRI Tech SolutionsRole: AI/ML EngineerLocation: Charlotte, NC (Hybrid - 3 Days onsite/week)Note: We need profiles who are local to... ...Charlotte, NC.Job Description: 6+ years in AI/ML engineering, data science, or advanced analytics.Hands-on experience...Full timeLocal areaRelocation3 days per week
- AI/ML Engineer**** Please note: This role is not eligible for 100% remote... ...and must be willing to be onsite at the client and/or Slalom office... .... We work across modern AI platforms, develop novel AI solutions,... ..., or AI.* Familiarity with leading data platforms like Databricks...Temporary workWork at officeLocal area3 days per week
- ...ability to manage small to medium projects. Candidates must have an Associate's Degree in Engineering and relevant experience in critical environments. The role is fully onsite with a focus on enhancing the customer experience through effective problem-solving and strategic...Full time
- ...collaborative team of architects, engineers, financial analysts, and delivery experts... ...with extensive experience in cloud/platform engineering, data/AI, or SRE who has spent years driving FinOps... ...scale. You have proven experience leading multi-workstream, multi-cloud...Full timeWork experience placementLive inWork at officeLocal area
$55 - $60 per hour
...immediately hiring for a Senior Site Reliability Engineer (SRE) Position Type: Full-time (Contract )... ...Engineer (SRE) you will: Lead infrastructure provisioning and... ...scalable microservices architectures. Platform Reliability & Cloud Engineering- Ensure...Hourly payFull timeContract workTemporary workWork experience placementImmediate startWorldwideFlexible hours- ...GorusuCompany: SRI Tech SolutionsJob Title: ML EngineerLocation: Charlotte , NC (hybrid... ..., R, etc.)Experienced in using AI/ML platforms, technologies, techniques (e.g.... ...machine learning algorithmsPerform feature engineering, model selection, and hyperparameter tuning...Hourly pay2 days per week
- ...reusable customizations for non-ML, ML, and deep learning... ...secrets management)Observability & SRE• Prometheus/Grafana, logging,... ...in ML model development, data engineering, and software engineering principles... ...Kubernetes-based serving platforms: o KServe, Kubernetes ML Serving...Full timeTemporary workRelocation
- ...building, securing, and operating VPI's cloud platform and engineering infrastructure. The individual will... ...supporting VPI's investment platform.Lead AWS platform architecture,... ...observability platforms, monitoring solutions, and SRE practices.Familiarity with financial services...Full time
- ...Cognitive Linguist Seeking a Cognitive Linguist - Onsite for a contract position located in Charlotte, NC, Pennington, NJ, or Plano... ...end-user intent across a multi-channel Virtual Assistant platform. This role focuses on designing scalable conversational AI solutions...Contract work
- ...DescriptionCapTech Machine Learning Engineers are responsible for... ...enterprise-level data platforms (e.g., AWS, Azure, GCP)... ...).Productionizing ML systems with a focus on... ...events, and leading teams of junior data scientists... ...must be available to work onsite in a client location or...Work at officeRemote workVisa sponsorshipWork visaFlexible hours
$119k - $187k
About this role:Wells Fargo is seeking a Lead Software Engineer in CIO Digital Innovation as a part of technology. Learn more about the career... ...DeliveryCollaborate with Product Owners, Architects, QA teams, SRE teams, and business stakeholders.Participate in Agile...Full timeWork experience placement- ...PFB the JD must have skills Hands on SRE Engineer with good analytical skills Good Exposure to both incident and Problem Management... ...(Servers,Load Balancer/Trace Logs) and have the acumen to lead any troubleshooting efforts from end to end until permanent solution...Permanent employmentFull time
- ...skilled and experienced Machine Learning Engineer to design and implement solutions for extracting... ...clustering. Proficiency in Python and ML libraries (TensorFlow, PyTorch, Hugging... ...• Expertise in APIs, REST, and cloud platforms (AWS, Azure, GCP). • Knowledge of distributed...
- ...Job Title: Machine Learning Engineer (GCP, Vertex AI, Dataproc, Apache Iceberg) Location: Onsite in Charlotte, NC Duration:... ...learning solutions on Google Cloud Platform (GCP). The successful... ...Iceberg to create production-grade ML pipelines capable of processing...
- ...ML Engineer Charlotte, North Carolina, United States Or refer someone Job Openings ML Engineer About the Job Our client is a rapidly... ...customer experiences. About Us Catalyst Labs is a leading talent agency with a specialized vertical in Applied AI, Machine...Full time
- ...Title: Data Center Infrastructure Engineer Location: Charlotte, NC Hybrid - 3 days onsite/week| Onsite Duration: 6-12+ Months Summary: ~ Client is seeking a Data Center Infrastructure Engineer for a hybrid contract-to-hire opportunity with a global...Contract work3 days per week
- AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting... ...partnerships with over 400 leading technology providers, our 10,0... ...seeking an experienced AI/ML Engineer to design, deploy, and operate... ...cloud-native machine learning platforms. You will work within agile,...Full timeContract workLocal areaRemote workFlexible hours
$122k - $200k
...!Position Summary:The Senior Software Engineer - Kubernetes Platform Engineering, Service Mesh & Developer... ...lifecycle management.The individual will lead the architecture and engineering of... ...with architecture, security, networking, SRE, DevOps, and application teams to...Full timeWork at officeFlexible hoursShift workDay shift- ...concepts in simple, understandable language· Ability to mentor and lead a team of consultants· Ability to work under minimal supervision... ...- follow timelines, schedule meetings, define agenda, guide onsite and offshore, define solution, get it validated and sign-off. Define...Permanent employmentFull timeH1b
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead ML Platform Engineer (SRE / FTE / Onsite). Be the first to apply!
- lead infrastructure engineer Charlotte, NC
- lead network engineer Charlotte, NC
- lead engineer Charlotte, NC
- lead operating engineer Charlotte, NC
- machine learning ai engineer Charlotte, NC
- computer vision machine learning engineer Charlotte, NC
- machine learning engineer Charlotte, NC
- ai ml engineer Charlotte, NC
- machine learning software engineer Charlotte, NC
- senior platform engineer Charlotte, NC


