Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

$207k - $300k

Google

Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.Design and implement inference optimization techniques.Investigate and resolve complex model inference performance bottlenecks across the stack.Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.Minimum qualifications:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.8 years of experience in software development.Experience in Python and C++, including navigating, debugging, and modifying serving codebases.Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.Preferred qualifications:Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.Experience with observability and reliability for large distributed systems.Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.8 years of experience in software development.Experience in Python and C++, including navigating, debugging, and modifying serving codebases.Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind in Mountain View, CA vacancy
  • $207k - $300k

    Write low-level performance software for a new systolic array accelerators...  ...performance software engineers and general software...  ...functional flows.At DeepMind our mission is to...  ...within DeepMind’s GenAI Systems, operating...  ...compilation, performance optimization, libraries, operating... 
    Performance

    Google

    Mountain View, CA
    2 days ago
  • $207k - $301k

     ...and tools for GenAI developers and...  ...agent quality performance at scale, design...  ...experience in software development.5...  ...ML design and optimizing ML...  ...degree or PhD in Engineering, Computer Science...  ...closely with Google DeepMind Research teams...  ...Cloud. As a Staff Software Engineer... 
    Performance

    Google

    Sunnyvale, CA
    12 hours ago
  •  ...THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications...  ...team. This role focuses on improving performance, efficiency, and scalability of...  ...large-scale models, working across the software-hardware stack.THE PERSONThe ideal... 
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $262k - $364k

     ...efficiently, prioritizing both performance and seamless user...  ...best practices and optimization in the day to day...  ...distributed team of engineers.Minimum qualifications...  ...years of experience in software development.5 years...  ...inventions. At Google DeepMind, we are a pioneering... 
    Performance

    Google

    Mountain View, CA
    1 day ago
  • $228k - $285k

     ...Staff Software Engineer, ML Training And Inference InfrastructureRivian is on a mission to keep the world adventurous forever. This goes...  ...of large autonomous driving models; and optimizing the training and inference performance.Responsibilities:Optimize the performance... 
    Performance
    Full time

    Rivian

    Palo Alto, CA
    4 days ago
  • $207k - $300k

     ...execution of a high-tempo engineering team, unblocking...  ...the timely launch of performant, high-quality features...  ...of experience in software development.5 years of...  ...architecture, performance optimization, benchmarking,...  ...algorithms.At Google DeepMind our mission is to build... 
    Performance

    Google

    Mountain View, CA
    3 days ago
  •  ...leading training and inference speeds; over 10 times...  ...the RoleWe're hiring a Staff Engineer to own major areas of...  ...bursty AI workloads, performance at high QPS, and the...  ...of experience in software engineering, with substantial...  ...at scale.Experience optimizing latency, throughput,... 
    Performance

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $193.93k - $352.29k

     ...pipelines to building to deploying the optimized models on Nuro’s fleet of self-driving...  ...road validation.Maintain an in-house ML inference platform to serve large language models...  ...excellence: Develop with a high standard for performance, scalability, and code quality.Domain... 
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  •  ...Sunnyvale, CA is seeking a Member of Technical Staff (Software Engineer) to implement infrastructure for high-performance, low-latency inference services. Applicants should have a...  ...involves deploying Kubernetes services, optimizing resource allocation, and collaborating... 
    Performance

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • Cerebras Systems, Inc. is seeking a Software Engineer in Sunnyvale, California to enhance high-performance, low-latency inference infrastructure. This role involves deploying scalable services, optimizing resource allocation, and integrating with containerized environments... 
    Performance

    Cerebras Systems, Inc.

    Sunnyvale, CA
    1 day ago
  • $174k - $252k

     ...physical synthesis tool runs.Engineer modular carve-outs of...  ....Build and optimize hardware compilation pipelines...  ...of experience with software development in C++.3 years...  ...design Power, Performance, and Area (PPA) closure...  ...transformative inventions. At DeepMind, we are a pioneering... 
    Performance

    Google

    Mountain View, CA
    2 days ago
  • $262k - $365k

     ...modeling, tuning, and optimization techniques to...  ...model (LLM) performance, quality, and capabilities...  ...Research and DeepMind to advance LLM...  ...of experience in software development.7 years...  ...state of the art GenAI techniques (e.g.,...  ...degree or PhD in Engineering, Computer Science... 
    Performance

    Google

    Mountain View, CA
    12 hours ago
  • Cerebras Systems is seeking a Software Engineer to build and maintain high-performance, low-latency inference infrastructure. You will deploy scalable inference services, optimize auto-scaling, and integrate with Docker/Kubernetes in production environments. You will collaborate... 
    Performance

    Cerebras

    Sunnyvale, CA
    2 days ago
  • $198k - $326k

     ...centered on trust and optimized for culture,...  ...meaning it will be performed both from home and...  ...recommendations, search, ads, GenAI, and other AI-...  ..., and enforceable engineering capabilities.As a Sr. Staff Software Engineer, you will...  ..., validation, inference, or deployment.... 
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    8 hours ago
  •  ...career. TTHE ROLE: Senior level engineer who will be responsible for...  ...’s strategy, architecture, optimization and tooling to achieve...  ...Pre-training and Distributed Inference Performance on AMD GPU. You will...  ...attainment across the entire software stack focusing on getting the... 
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $179.5k - $260k

     ...re looking for an Applied AI Engineer with strong backend and AI...  ...architect, build, and scale secure, performant systems. You'll work closely...  .... Integrate LLMs and GenAI components into production workflows...  ...to productionize models, optimize inference pipelines, and collect... 
    Performance
    Full time
    Flexible hours
    Night shift

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $174k - $252k

    Design and develop robust software to integrate hardware components...  ...Computer Science, Computer Engineering, Robotics, or a related...  ...gate array (FPGAs), and high-performance computing platforms.Proficiency...  ...engineering teams.At Google DeepMind our mission is to build the... 
    Performance

    Google

    Mountain View, CA
    2 days ago
  • $262k - $365k

     ...maintain, and enhance large scale software solutions.Provide technical...  ...coach a distributed team of engineers.Drive technical project...  ...large-scale ML infrastructure optimization, and oversee the design and implementation...  ...of state-of-the-art GenAI solutions.Minimum... 

    Google

    Sunnyvale, CA
    12 hours ago
  • $277k - $308k

     ...experience in either technical product/engineering/researcher roles, or policy/...  ...specialist. About The Job The GenAI Technical Engagement team within Google DeepMind is responsible for translating...  ...insights for campaign optimization. Google is proud to be an equal... 

    Google DeepMind

    Mountain View, CA
    4 days ago
  • $207k - $300k

     ...architect high-performance model...  ...system efficiency.Optimize Infrastructure...  ...Manage and Mentor Engineers: Lead and grow...  ...team of ML and software engineers,...  ...state of the art GenAI techniques (e....  ...groups like Google DeepMind, Google...  ...advertising.As a Staff Machine... 
    Performance

    Google

    Mountain View, CA
    8 hours ago
  • $174k - $253k

    Senior Software Engineer, Gemini Audio, DeepMind Location: Mountain View, CA, USA; New York...  ...with compiler optimization, code generation, and runtime...  ...Learning Optimization, Performance Optimization, and Large...  ...the model training and inference efficiency. Information... 
    Performance

    Google

    Mountain View, CA
    4 days ago
  • $189k - $274k

     ...accessible for all. As a Staff Software Engineer focusing on Deep Learning...  ...pivotal role in enhancing the performance of Deep Learning networks...  ...performance analysis and optimization of these networks, ensuring...  ...Triton, Mojo and other inference acceleration tools.For roles... 
    Performance
    Work at office
    Local area
    3 days per week

    Aurora Innovation

    Mountain View, CA
    1 day ago
  • $262k - $365k

     ...Android, Chrome, and more. Improve performance of on-device model inference via optimizations in the model representation, on...  ....8 years of experience in software development.7 years of experience...  ...:Master’s degree or PhD in Engineering, Computer Science, or a related... 
    Performance

    Google

    Sunnyvale, CA
    12 hours ago
  • $210k - $250k

     ...OpportunityWe're seeking a talented Staff Engineer to join our Frontend Team....  ...leadership on enhancing performance, infrastructure, and overall...  ...in writing clean, reusable, optimized HTML/CSS, Typescript and...  ...early investors in Google, DeepMind, Zoom, and Tesla.Otter.ai is... 
    Performance
    Permanent employment

    Otter.ai

    Mountain View, CA
    2 days ago
  • $189.3k - $290.7k

     ...Summary We are seeking a Staff Software Engineer to join our ADAS HMI team...  ...on-device perception and inference, driver-state interpretation...  ...technical trade-offs involving performance, memory, compute, startup...  ...Experience diagnosing and optimizing complex integrated systems... 
    Performance
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    8 hours ago
  • $229k - $343k

     ...services.We’re looking for a Staff Software Engineer to join Snap Inc on our...  ...ensuring they meet correctness, performance, and reliability...  ...reliably for large-scale batch inference and low-latency online...  ...architectureProficiency with performance optimization techniquesExperience... 
    Performance
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    12 hours ago
  • Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and...  ...will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source... 
    Performance

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    2 days ago
  • $198k - $326k

     ...centered on trust and optimized for culture,...  ...meaning it will be performed both from home...  ...training, feature engineering and serving with...  ...data infra, compute software, and hardware to...  ...enable GPU based inference for a large...  ...at scale.As a Sr. Staff Software Engineer... 
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    2 days ago
  • Decisive Point is seeking a Software Engineer in Sunnyvale, California, with expertise in optimizing machine learning models for embedded systems. This role involves performance optimization for embedded compute platforms, collaborating with ML engineers, and requires strong... 
    Performance

    Decisive Point

    Sunnyvale, CA
    1 day ago
  • Applied Intuition Inc. is seeking a software engineer with deep expertise in optimizing ML models for production-grade embedded runtime environments. You...  ...engineers and software developers, profiling model performance, implementing pruning/quantization, and driving efficient... 
    Performance

    Applied Intuition Inc.

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind. Be the first to apply!