Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineering Manager

GoFundMe

Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations) In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe’s machine learning and AI systems reliable, scalable, and safe in production This role requires strong technical judgment across the ML lifecycle (data → training → online inference → monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents Drive operational strategy for GoFundMe’s generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability Collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy stakeholders to translate business goals into team priorities and measurable operational outcomes Manage vendor and platform relationships (e.g., cloud ML platforms, LLM providers) and make build-vs-buy calls that balance cost, control, and speed Report on team health, system reliability metrics, and operational risk to senior engineering leadership Employ a diverse set of tools and platforms, including Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, and GitHub, to guide your team in developing, deploying, and maintaining scalable and robust machine learning systems Benefits $600 annual fitness and wellness reimbursement Wide range of health insurance options, including medical, dental, and vision (GoFundMe covers 100% of employee premiums, and 80% of spouse and dependents) Weekly massages Standing desks Fully-stocked kitchens & daily lunches Team off-sites & monthly social events Many of our offices are dog friendly Enhanced parental leaves 10 paid holidays, 17 days of accrued vacation per year, unlimited sick time & three volunteer days Caltrain GoPasses for our Bay Area commuters $50/month for employees commuting to and from work (public transit and/or parking) Quarterly volunteer events in each office to give back to our local communities “Gives Back” program, where employees nominate fundraisers weekly for donations from GoFundMe 401(k) retirement plan with company matching Access to learning tools and resources, including a subscription to Udemy, guest speakers, and internal brown bag sessions 1-3+ years of experience directly managing engineers, ideally in an MLOps, ML platform, or infrastructure context, with a track record of hiring and developing strong teams Familiarity with generative AI/LLM infrastructure and operational considerations (latency, cost, safety guardrails) is a strong plus Advanced degree (Master’s or Ph.D.) in Computer Science, Statistics, Data Science, or a related technical field is preferred Proven experience implementing ML monitoring for both technical and business metrics (drift, calibration, segment performance, latency, error budgets) and running models reliably in production Sense of humor is optional but appreciated Experience designing and operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies (canaries, shadowing, backward-compatible schema evolution) Strong proficiency in Python and ML libraries/frameworks such as PyTorch, TensorFlow, Scikit-learn, plus strong software engineering fundamentals (testing, code review, CI/CD, API design, performance, and reliability) — enough depth to stay hands-on and credible with your team Strong data engineering fluency: building reliable datasets and features using SQL, Spark/Databricks, and warehouse technologies (e.g., Snowflake), with an understanding of event semantics, identity resolution, and data quality controls Strong leadership and mentoring skills and a proven ability to raise the bar on architecture, engineering quality, and operational rigor for production ML/AI systems 7+ years of hands-on experience building and shipping production machine learning systems, with demonstrated ownership of backend services and ML pipelines in a high-availability environment Ability to break down ambiguous, high-impact problems, define crisp interfaces and success metrics, and deliver iteratively while managing stakeholder expectations across engineering leadership, product, and data science #J-18808-Ljbffr GoFundMe

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineering Manager. Be the first to apply!