Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Machine Learning Engineer - LLM Quantization & Deployment

$174.72k - $295.68k

XPENG Motors

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.Our mission is to build strong foundation for LLM deployment and quality sign-off for next-gen XPENG Turing AI chip. This includes and is not limited to: LLM model fine tuning, PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling.Build robust model export, calibration, benchmarking, validation, and deployment pipelines.Engage early with the VLA model research team to establish performance estimates and prove model feasibility.Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance.Analyze numerical errors, accuracy regressions, and performance trade-offs.Develop PTQ and QAT orchestration workflows.Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.Collaborate with the in-vehicle software team on latency analysis and issue triage.Collaborate with the training infrastructure team to develop QAT and model distillation.Basic QualificationsMaster in CS/CE/EE, or equivalent, with 1-3 years of industry experience. Open to new graduates.Strong understanding of Transformer architectures and LLM inference.Hands-on experience quantizing or deploying deep learning models in production.Proficiency with PyTorch and at least one inference or compilation stack.Strong Python programming and software engineering skills.Ability to work effectively across research, systems, infrastructure, and product teams.Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.Preferred QualificationsExperience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization.Experience with AWQ, GPTQ, SmoothQuant, or related methods.Strong numerical analysis and systems engineering skills.Experience with one or more LLM runtimes, such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes.Experience deploying LLMs on resource-constrained or heterogeneous hardware.Contributions to model optimization, inference, compiler, or serving projects.Publications at NeurIPS, ICML, ICLR, ACL, or related conferences.What We ProvideA fun, supportive and engaging environment.Infrastructures and computational resources to support your work.Opportunity to work on cutting edge technologies with the top talents in the field.Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.Competitive compensation package.Snacks, lunches, dinners, and fun activities.The base salary range for this full-time position is $174,720 - $295,680, in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.We are an Equal Opportunity Employer. It is our policy to provide equal employment opportunities to all qualified persons without regard to race, age, color, sex, sexual orientation, religion, national origin, disability, veteran status or marital status or any other prescribed category set forth in federal or state regulations.

Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Senior Machine Learning Engineer - LLM Quantization & Deployment in Santa Clara, CA vacancy
  • Research, design, development, and deployment of advanced AI agents and agentic systems...  ...product managers, UX designers, and other engineers to define requirements and deliver...  ...command execution. Knowledge and passion in machine learning algorithms, GenAI, LLMs, and Agentic... 
    Senior
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    9 hours ago
  • $184.7k - $324.8k

    Senior Machine Learning Engineer Apple Services Engineering (ASE) builds experiences...  ...expertise in Generative AI, LLM architectures, and...  ...research inception to production deployment, and who thrives in environments...  ...optimization techniques (quantization, distillation, model... 
    Senior
    Relocation

    Apple

    Cupertino, CA
    3 days ago
  • $227k - $300k

    Senior Staff Machine Learning Engineer At Sonatus, we're driving the transformation to AI...  ...funding and proven by global deployment, we're solving some of...  ..., including cloud-based LLM APIs (Gemini, OpenAI, Claude...  ...or embedded NPUs. Apply quantization, pruning, distillation,... 
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    5 days ago
  • $224k - $356.5k

    NVIDIA is looking for a Machine Learning Engineer to join the GPU accelerated Apache Spark team.Apache...  ...with key partners and customers on the deployment of complex machine learning solutions...  ...sophisticated ML methodologies, including LLM/GenAI, reinforcement learning, and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $188.5k - $282.7k

     ...Semantic AI Governance Engine, which is the first...  ...At its core, SAGE is "LLM-as-judge" applied to...  ...training choices (LoRA, quantization-aware training,...  ...fleets.Optimizing live deployments through shared GPU pools...  ...in Computer Science, Machine Learning, Computer Engineering... 
    Senior
    Permanent employment
    Local area

    Rubrik

    Palo Alto, CA
    a month ago
  • $200k - $280k

    Machine Learning Engineer Responsibilities: Develop, optimize, and deploy lightweight machine learning models for edge AI applications, particularly for audio processing...  ...and model compression techniques, including quantization, pruning, and knowledge distillation.... 
    Senior
    Local area

    TetraMem - Accelerate The World

    San Jose, CA
    5 days ago
  • $153.75k - $225k

     ...operating at scale across Azure and AWS, deployed in multiple regions globally, including...  ...collaboration, and high standards. Our engineers, product leaders, and go-to-market...  ...Qualifications:Knowledge and passion in machine learning algorithms, GenAI, LLMs, and Agentic AIUnderstanding... 
    Senior
    Work experience placement
    Work at office
    3 days per week

    Eightfold

    Santa Clara, CA
    2 days ago
  • $220k - $300k

    Senior Principal Machine Learning Engineer San Jose, California, United States The era of pervasive AI has arrived. In this era, organizations...  ...the RDU. This critical role bridges advanced LLM research and practical deployment, involving the development of model... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova Systems

    San Jose, CA
    5 days ago
  • $174.72k - $295.68k

     ...cutting-edge R&D in AI, machine learning, and smart...  ...time Machine Learning Engineer / Research Scientist to...  ...to design, train, and deploy large-scale multi-modal...  ...optimization, including quantization, export, and latency-accuracy...  ...driving models, or LLM/VLM architectures (e.g... 
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $130k - $220k

     ...future of autonomy, Plus is looking for talented individuals to join its fast-growing teams. We’re looking for a machine learning engineer to train and deploy the latest generation of ML-based planning algorithms on the extensive data we collect every day across our... 
    Senior
    Full time

    Plusai

    Santa Clara, CA
    9 hours ago
  •  ...from there. Real enterprise deployments, not demos. Role Overview: As a Senior MLOps Engineer, you will own the infrastructure...  ...of ML engineering, LLM inference infrastructure, and...  ...inference‑time optimizations — quantization (AWQ, GPTQ, FP8/GGUF), distillation... 
    Senior
    Full time

    NACE

    Palo Alto, CA
    3 days ago
  •  ...performing backend systems for machine learning products. Build ML...  ...and retrieval. Build and deploy machine learning models across...  ...related field plus 3+ years of engineering experience and 5+ years of...  ...advertising industry experience, LLM-based data generation or... 
    Senior
    Full time
    Work experience placement

    Apple

    Cupertino, CA
    3 days ago
  • $140k - $210k

     ...make a difference at Fiserv.Job TitleSenior Machine Learning EngineerWhat does a successful Senior Machine Learning Engineer do at Clover?The Data Science team at Clover...  ...drive the architecture design, development, and deployment for all AI/ML products the Data Science team... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    6 days ago
  • $230k - $265k

     ...want to lead projects to build and deploy cutting-edge AI technology to help people...  ...industry-veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll bring your strong software...  ...large-scale SID / ASR / NLP / LLM systems that power mission-critical... 
    Senior
    Permanent employment

    Otter.ai

    Mountain View, CA
    a month ago
  •  ...Responsibilities Develop and improve machine learning and large language model systems for...  ...generation, debugging, and related software engineering tasks. Build AI-powered developer...  ..., model development, evaluation, deployment, and production iteration. Strong system... 
    Senior
    Full time

    ByteDance

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...systems at scale. We seek a Senior ML Engineer to compose and deliver...  ...proven experience in applied machine learning or AI research.Deep...  ...understanding of ML development and deployment life cycle with the...  ...experience with model compression, quantization, and real-time inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    18 days ago
  •  ...Bosch, and DSV are working with Plus to accelerate the deployment of next-generation autonomous trucks. If you’re ready to...  ...individuals to join its fast-growing teams. We are seeking a Senior Machine Learning Engineer with expertise in deep learning and data analysis. In... 
    Senior

    PlusAI

    Santa Clara, CA
    23 days ago
  • $148.75k - $361k

     ...scale and with low latency. We use Machine Learning, Reinforcement Learning, AI, Control...  ...seeking a talented and experienced Senior Software Engineer, MLOps/DevOps, to join the Advertising...  ...that accelerate ML experimentation and deployment at internet scale. You will partner... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    a month ago
  •  ...running. Our Team's Vision Our Engineering team is shaping the future of...  ...culture, come join us! As a Senior Software Engineer, you will...  ...managing complex multi-region cloud deployments. Vector DB Optimization:...  .... AI Ops: Experience with LLM deployment optimization (e.g.... 
    Senior
    Immediate start

    Illumio

    Sunnyvale, CA
    5 days ago
  • $202k

     ...efficiency no one else can match. As a Senior ML Engineer, you will be at the forefront of...  ...development and implementation of the latest machine learning techniques across the autonomy...  ...engineering teams to enable the successful deployment of the latest machine learning... 
    Senior
    Full time
    Work experience placement
    Work at office
    Remote work

    Uber Technologies Inc

    Sunnyvale, CA
    5 days ago
  • $160k - $200k

     ...with Plus to accelerate the deployment of next-generation...  ...fast-growing teams. As a Senior ML Infrastructure Engineer at Plus, you will design scalable...  ...state-of-the-art deep learning frameworks like PyTorch or...  ...boundaries of what's possible in machine learning infrastructure... 
    Senior

    PlusAI

    Santa Clara, CA
    9 days ago
  • Sr. Machine Learning Engineer It started with a simple idea: what if surgery could...  ...cancer. We are seeking a senior algorithm engineer to play...  ...and execute model deployment pipelines, seamlessly integrating...  ...utilizing techniques like quantization, pruning, and hardware... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive

    Sunnyvale, CA
    5 days ago
  •  ...working with Plus to accelerate the deployment of next-generation autonomous trucks...  ...Responsibilities Design, develop, and deploy learned behavior planning models using...  ...MS or PhD in Computer/Software Engineering, Robotics, Machine Learning or related field.... 
    Senior

    PlusAI

    Santa Clara, CA
    18 days ago
  • $151.8k - $265.35k

     ...and a media-intelligence layer, and deployed across new and existing Adobe...  ...adjacent verticals. We are hiring a Senior Machine Learning Engineer to build the pipelines and services...  ...optimization for latency and cost — quantization, batching, and serving runtimes; custom... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    1 day ago
  • $124k

     ...scale. In addition, we deploy these models to edge...  ...developing post-training quantization and quantization-aware...  ...making massive deep learning models run lightning-fast...  ...compiler, inference engine, and silicon teams to...  ...in Computer Science, Machine Learning, Robotics, Computer... 
    Hourly pay
    Full time
    Temporary work
    Immediate start
    Flexible hours

    Tesla

    Palo Alto, CA
    2 days ago
  • $232k - $310k

     ...scale across Azure and AWS, deployed in multiple regions globally...  ...collaboration, and high standards. Our engineers, product leaders, and go-to-...  ...the boundaries of applied machine learning. We work with massive...  ...(vLLM, TensorRT-LLM).Desired Skills & Experience... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    3 days ago
  • $170k - $185k

     ...that are governed, secure, always learning, and working 24/7 to drive...  ...Job Description As a Senior AI/ML Engineer, you will design, develop, and deploy enterprise-scale AI systems that...  ...with AI evaluation frameworks, LLM observability, AI governance, and... 
    Senior
    Full time
    For contractors
    Flexible hours

    IFS

    Palo Alto, CA
    23 days ago
  • $170.6k - $261.3k

     ...transportation on a global scale.As a Senior Machine Learning Engineer on the State Estimation and Mapping...  ...techniques (e.g., pruning, quantization, distillation) to meet on‑vehicle compute...  ...interfaces, configuration, deployment, monitoring, and regression safeguards... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Machine Learning Applications and Compiler Engineer!NVIDIA is seeking engineers to develop algorithms and optimizations for our...  ...libraries, tooling, and interfaces that enable seamless deployment of models across platforms.Benchmark, profile, and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $151.8k - $265.35k

    Senior AI / Machine Learning Engineer — Fraud DetectionThe OpportunityWe are in search of an experienced Senior AI/ML Engineer to develop and broaden...  ...to end — from signals and modeling through production deployment and real-time decisioning. Join us in crafting best in... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Worldwide
    3 days per week

    Adobe Systems

    San Jose, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Machine Learning Engineer - LLM Quantization & Deployment. Be the first to apply!