Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, ML Inference Performance

Full-time

Sambanova Systems


The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

About The Role


The Principal Compiler Engineer - ML Systems position will be responsible for working with the different layers of the compiler stack and coordinating with other development teams here at SambaNova. It is a critical role responsible for driving innovation in compiler infrastructure and optimization algorithms that enable state-of-the-art ML model performance on the SambaNova platform. This can involve anything from digging through PyTorch and machine learning models to determining how to map operations on to our underlying hardware.

Responsibilities



  • Lead compiler engineering through ensuring standard methodologies, enterprise product insertion and process evolution.

  • Work with peers, domain experts, developers, customers, and work across the enterprise seeking optimal solutions.

  • Develop, integrate, and implement products.

  • Provide support for proposals in key areas aligned with core team competencies.

Basic Qualifications



  • Bachelor’s or Master’s Degree in Computer Science, Computer Engineering, or equivalent with 5-10 years of industry experience.

Additional Qualifications



  • Deep theoretical understanding of compiler fundamentals.

  • Experience building and deploying software products.

  • Experience with one or more deep learning frameworks (i.e. TensorFlow, PyTorch) is a plus.

  • Experience with common compiler development practices and methodologies.

  • Excitement about high-performance systems engineering and performance debugging.

  • An appreciation for process and developing cross-disciplinary collaboration.

Preferred Qualifications



  • Experience with MLIR.

  • Familiarity with machine learning models and frameworks.

  • Familiarity with accelerated computing.

  • Exposure to dataflow architectures.

 


Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Vacancy posted 20 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, ML Inference Performance in Palo Alto, CA vacancy
  • $160.36k - $240.54k

     ...other leading investors.About the RoleThe ML Infrastructure team is responsible for...  ...road validation.Maintain an in-house ML inference platform to serve large language models...  ...excellence: Develop with a high standard for performance, scalability, and code quality.Domain... 
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $160.36k - $240.54k

     ...leading investors. About the Role The ML Infrastructure team is responsible for...  ...validation. Maintain an in-house ML inference platform to serve large language models...  ...excellence: Develop with a high standard for performance, scalability, and code quality.... 
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact...  ...of TensorRT, NVIDIA’s SDK for high-performance deep learning inference.Closely...  ...TensorFlow, ONNX Runtime or other ML frameworks.NVIDIA is widely considered... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    20 hours ago
  • $228k - $285k

     ...protect it for future generations. Role SummaryAs a Staff Software Engineer, ML training and inference infrastructure, you will be a member of the...  ...driving models; and optimizing the training and inference performance. ResponsibilitiesDesign, train, and deploy large deep... 
    Performance
    Full time
    Contract work
    Local area

    Rivian Automotive

    Palo Alto, CA
    4 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud...  ...hardware.As a Software Development Engineer on the Inference Model Enablement team, you will...  ...Key job responsibilities* Deliver high-performance models using distributed inference libraries... 
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $193.3k - $261.5k

     ...AWS) builds AWS Neuron, the software development kit used to...  ...s Inferentia and Trainium ML accelerators. This comprehensive...  ...enabling unparalleled ML inference and training performance.The Inference Enablement...  ...-software boundary, our engineers build systematic infrastructure... 
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    20 hours ago
  •  ...Luma Model Serving Engineer You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines...  ...RoCE, InfiniBand, NVLink). High-performance large-scale ML systems (100+ GPUs). CUDA, and... 
    Performance

    Luma AI

    Redwood City, CA
    2 days ago
  • $92k - $135k

     ...combines superior infrastructure performance with deep technical expertise...  ...What You'll Do: Join the Inference team to ship production...  ...mentorship from experienced engineers. About the role: Implement...  ...that deployed a microservice or ML inference demo. Coursework... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    17 days ago
  • $184k - $287.5k

     ...seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...’ll architect and implement high-performance inference stacks, optimize GPU...  ...pareto frontier for the field of ML Systems; survey recent publications... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    20 hours ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...the Role We're hiring a Software Engineer to help contribute to...  ...technical areas. Reliability & Performance. Architect active-active systems...  ...Influence. Partner with ML, Product, Infrastructure, and... 
    Performance
    Full time

    Cerebras Systems

    Sunnyvale, CA
    20 hours ago
  •  ...We are at the forefront of software and hardware innovation, pushing...  ...Principal System Software Engineer, AI Inference ExecutionWhat you will do:...  ...with other software (ML and compilers) and hardware...  ...toolsExperience with distributed, high-performance software design and... 
    Performance
    3 days per week

    d-Matrix

    Santa Clara, CA
    20 hours ago
  • $153k - $222k

     ...defense. As a full-stack engineer, you'll be involved in all...  ...bringing the latest and greatest software advancements to the...  ...learning model training, and inference Work on ML-adjacent tooling and infrastructure...  ...as it grows, ensuring performance, reliability, and security... 
    Performance
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Mountain View, CA
    20 hours ago
  • $174k - $253k

     ...improve the model training and inference efficiency.Minimum...  ...experience.5 years of experience with software development in Python.3 years...  ...Machine Learning Optimization, Performance Optimization, and Large...  ...Experience tailoring algorithms and ML models to exploit TPU... 
    Performance

    Google

    Mountain View, CA
    4 days ago
  • $152k - $241.5k

     ...We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior...  ...PyTorch, JAX.Knowledge of close-to-metal performance analysis, optimization techniques, and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.93k - $291.15k

     .... We are looking for strong software engineers to research, develop, and implement...  ...generalizable and scalable ML planner that can power L4...  ...platforms in a safe, performant, and scalable way.Provide technical...  ..., distributed training, or inference optimization.At Nuro, we... 
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    20 hours ago
  • $213.51k - $245k

     ...their careers. We’re a high-performing, fast-moving team with ethics...  ...for an exceptional Senior Software Engineer to help shape the future of...  ...implementation of end-to-end ML pipelines, including data ingestion...  ..., ensuring low-latency inference and reliable service... 
    Performance
    Work at office
    Remote work
    Flexible hours
    Shift work
    3 days per week

    Robinhood Financial

    Menlo Park, CA
    2 days ago
  • $193.3k - $261.5k

    Build the large-scale ML training and real-time inference systems that deliver highly...  ...If you're energized by ML engineering at massive scale with direct...  ..., optimize model performance at scale, and implement end...  ...-internship professional software development experience- 5... 
    Performance
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Palo Alto, CA
    1 day ago
  • $165.2k - $223.6k

     ...talented scientists and engineers to innovate on behalf...  ...professional software development experience...  ...architecture, training/inference lifecycles, and optimization...  ...qualification - Knowledge of ML frameworks including...  ...- Knowledge of system performance, memory management,... 
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    3 days ago
  • $119.8k - $234.7k

     ...Less than 25%Profession: Software EngineeringDiscipline:...  ...as a Service, Azure ML, Cognitive Services,...  ...a Principal Software Engineer - Responsible AI who is...  ...implementation and with high performance, low latency, and high...  ...AI models and agents Inference, routing,... 
    Performance
    Ongoing contract
    Work at office
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    1 day ago
  • $180k - $258.75k

     ...Models.We are looking for a Senior Software Engineer to join our end-to-end automated...  ...in C++ and Python, that supports ML training, evaluation, and inference workflows.Build and maintain ML...  ...packaging, runtime integration, and performance validation on embedded compute... 
    Performance
    Full time
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    4 days ago
  • $160.36k - $240.54k

     ...leading investors. About the Role The ML Infrastructure team is responsible for...  ...validation. - Maintain an in-house ML inference platform to serve large language models...  ...excellence: Develop with a high standard for performance, scalability, and code quality. -... 
    Performance
    Full time
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    11 hours ago
  • $193.93k - $352.29k

     ...autonomously inside Nuro's own engineering organization, under...  ....About You3+ years of software engineering experience...  ...under the hood at inference. Attention and KV-...  ...helped.Experience with ML training or research infrastructure...  ...for an annual performance bonus, equity, and a... 
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $157k - $235k

     ...and other digital services.Snap Engineering teams build fun and technically sophisticated...  ...forefront.We’re looking for a Software Engineer to join the ML Platform Experience team, part of...  ...data generationDevelop high-performance inference systems to ensure fast and... 
    Performance
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    1 day ago
  • $150k - $195k

     ...Software Engineer DeepInfra is looking for early-career Software Engineers...  ...If you're excited about AI/ML, have built and shipped...  ...Design, develop, and test inference solutions for state-of-the-art...  ...new features, improve system performance, and contribute to overall system... 
    Performance
    Full time

    DeepInfra

    Palo Alto, CA
    4 days ago
  • $140k - $150k

     ...Software Engineer, Early Career DeepInfra is looking for early-career...  ...If you're excited about AI/ML, have taken related courses...  ...to design, develop, and test inference solutions for state-of-the-art...  ...experiment with improving model performance. Try new things. Ship... 
    Performance
    Full time
    Internship

    DeepInfra

    Palo Alto, CA
    4 days ago
  • $210.3k - $273.4k

     ...interacting with content, we're engineering the next generation of...  ..., self-driven Senior Software Engineer to raise the bar...  ...reliability, scalability, performance, and maintainability of our...  ...powering asset generation, LLM inference, and the ML-driven tools that bring AI... 
    Performance
    Temporary work
    Work at office
    Worldwide
    Relocation package

    Unity Technologies

    Mountain View, CA
    2 days ago
  • $145k - $200k

     ...builds the world’s leading software for data-driven...  ...Role We are a software engineering team with expertise in enabling ML models in production. We...  ...across the full stack, from inference engines, GPU scheduling...  ...ResponsibilitiesBuilding high-performance model serving... 
    Performance
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Palo Alto, CA
    1 day ago
  •  ...About the Role As a software engineering intern, you will work closely...  ...Platform, Onboard Systems, ML Infrastructure, Simulation,...  ...provide a reliable and high-performance platform that allows our autonomy...  ...-cloud training and onboard inference. Our solutions include a distributed... 
    Performance
    Internship

    Nuro

    Mountain View, CA
    20 hours ago
  • $150k

     ...highly motivated, and focused on engineering excellence. This organization...  ...to measure and improve performance. Work closely with product...  ...clean, efficient code for AI/ML systems. Hands-on experience...  ...scale distributed training and inference systems on Kubernetes.... 
    Performance
    Temporary work

    SpaceXAI

    Palo Alto, CA
    7 days ago
  • $188k - $275k

     ...combines superior infrastructure performance with deep technical...  ...more at What You'll Do: Inference Platform Team The Inference...  ...About the role: As a Staff Software Engineer (IC5) on the Inference team,...  ...Exposure to large-scale AI/ML infrastructure or hyperscale... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, ML Inference Performance. Be the first to apply!