Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - Model Performance

$220k - $320k
Full-time

Inference Corp

Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet you.

About Inference.net

Inference.net trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match GPT-5 accuracy but are smaller, faster, and up to 90% cheaper. Our platform handles everything end-to-end: distillation, training, evaluation, and planet-scale hosting.

We are a well-funded ten-person team of engineers who work in-person in downtown San Francisco on difficult, high-impact engineering problems. Everyone on the team has been writing code for over 10 years, and has founded and run their own software companies. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do. Most of us are in the office 4 days a week in SF; hybrid works for Bay Area candidates.

About the Role

You will be responsible for making our inference stack as fast and efficient as possible. Your work spans from implementing known optimization techniques to experimenting with novel approaches, always with the goal of serving models faster and cheaper at scale.

Your north star is inference performance: latency, throughput, cost efficiency, and how quickly we can bring new model architectures into production. You'll work across the full inference stack—from CUDA kernels to serving frameworks—to find and eliminate bottlenecks. This role reports directly to the founding team. You'll have the autonomy, a large compute budget, and technical support to push the limits of what's possible in model serving.

Key Responsibilities

  • Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving

  • Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance

  • Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure

  • Add support for new model architectures, ensuring they meet our performance standards before going to production

  • Experiment with novel inference techniques and bring successful approaches into production

  • Build tooling and benchmarks to measure and track inference performance across our fleet

  • Collaborate with applied ML engineers to ensure trained models can be served efficiently

Requirements

  • 2+ years of experience in ML systems, inference optimization, or GPU programming

  • Strong proficiency in Python and familiarity with C++

  • Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)

  • Deep understanding of GPU architecture and experience profiling GPU workloads

  • Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)

  • Experience with PyTorch and understanding of how models execute on hardware

  • Track record of measurably improving system performance

Nice-to-Have

  • Experience with CUDA programming

  • Familiarity with serving non-LLM models (TTS, vision, embeddings)

  • Experience with distributed inference and multi-GPU serving

  • Contributions to open-source inference frameworks

  • Experience with Docker and Kubernetes

You don't need to tick every box. Curiosity and the ability to learn quickly matter more.

Compensation

We offer competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.

Equal Opportunity

Inference.net is an equal opportunity employer. We welcome applicants from all backgrounds and don't discriminate based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status.

If you're excited about making AI inference faster for everyone, we'd love to hear from you. Please send your resume and GitHub to View email address on jobs.jobcopilot.com and/or apply here on Ashby.

Vacancy posted 18 hours ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - Model Performance in San Francisco, CA vacancy
  • $172.43k - $230.95k

     ...and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About This Role: The Senior Software Engineer for the AI Model Lifecycle team will play a crucial role in building a... 
    Senior
    Performance
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    18 hours ago
  • $166k - $225k

     ...to improve their business. Databricks’ Model Serving product provides enterprises with...  ...strong SLAs and cost efficiency.As a Senior Engineer, you’ll play a critical role in shaping...  ...architectural decisions and trade-offs to optimize performance, throughput, autoscaling, and... 
    Senior
    Performance
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    5 days ago
  •  ...at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently...  .... Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...access our start-of-the-art AI models, allowing them to do things...  ...able to before. We focus on performant and efficient model inference...  ...Role We are looking for an engineer who wants to take the world's...  ...least 5 years of professional software engineering experience.... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...next generation of frontier models. By co-designing chips, systems...  ...runtime within the inference engine that executes complex,...  ...layers of the cluster serving software stack, translating demanding...  ...silicon can deliver meaningful performance in production. In this role... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $124k - $280k

     ...OpportunityAs a Strategy& Strategy Consulting - Business Model Reinvention - Senior Manager you will provide strategic guidance and insights...  ...organizations, analyzing market trends and assessing business performance to develop recommendations that help clients achieve... 
    Senior
    Performance
    Full time
    H1b

    PwC

    San Francisco, CA
    3 days ago
  • $77k - $202k

     ...OpportunityAs a Strategy& - Strategy Consulting Business Model Reinvention - Senior Associate, you will provide strategic guidance and insights...  ..., analyzing market trends and assessing business performance to develop recommendations that help clients achieve their... 
    Senior
    Performance
    Full time
    H1b

    PwC

    San Francisco, CA
    3 days ago
  • $115.7k - $119.5k

     ...purpose. We work in a uniquely collaborative model across the firm and throughout all...  ...our clients to thrive.What You'll DoAs a Senior Analyst (SA) in a Client Focused role inside...  ...base salary, annual discretionary performance bonus, retirement contribution, and a market... 
    Senior
    Performance
    Work at office
    Local area

    The Boston Consulting Group

    San Francisco, CA
    4 days ago
  • $119k - $218.3k

    Position Summary Senior Consultant - Digital Assets Enterprise Strategy, Risk and Operating Model Design Enterprise Operations & Risk Ready for a fast-paced exciting...  ..., governance, processes, roles, skills, and performance measures.Build and validate business cases... 
    Senior
    Performance
    Contract work
    Work at office
    Local area

    Deloitte

    San Francisco, CA
    4 days ago
  • $298k - $368k

     ...diverse set of sensors, enabling engineers like you to (1) develop...  ...real-world data, to (2) develop models and model training at scale,...  ...architectures. Optimize model performance for on-device use cases (...  ...Engage directly with research, software engineering, hardware... 
    Senior
    Performance
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $325k

     ...company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production...  ...ideal candidate has over 5 years of software engineering experience, strong familiarity...  ...with researchers and focus on performance optimization. Compensation ranges... 
    Senior
    Performance

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...About the Role As a Senior Software Engineer on the Handshake AI Enterprise team, you'll own the...  ...TypeScript, with deep knowledge of frontend performance, component architecture, and design...  ...at a company with a forward-deployed or embedded engineering model... 
    Senior
    Performance
    Full time
    Work at office

    Handshake

    San Francisco, CA
    18 hours ago
  •  ...Opportunity Gradient is looking for a Founding Software Engineer to own how the robots move data from AI models to actuators. You will work close to the...  ...is the only stochastic component. You will ship performance-critical code to real robots daily.   What You... 
    Senior
    Performance
    Full time

    Rethink recruit

    San Francisco, CA
    1 day ago
  •  ...risk inherent to generative models. CTGT is the deterministic governance...  ...on HaluEval, the CTGT Policy Engine (paired with GPT-120B OSS)...  ...reliable, controllable, and performant in practice. Our mission...  ...is the fundamentals of software engineering, applied at the level... 
    Senior
    Performance
    Full time

    Ctgt

    San Francisco, CA
    1 day ago
  • PLEASE CLICK HERE TO SEE *ALL* OF OUR JOB OPENINGS!Senior Software Engineer — Backend PerformanceAs a Senior Software Engineer on Backend Performance, you own the hottest paths in data products — the code that has to be fast because everything downstream depends on it.... 
    Senior
    Performance

    Three Pillars Recruiting

    San Francisco, CA
    5 days ago
  •  ...efficient at scale. Role Overview We are seeking a Senior Software Engineer to design, build, and own core backend and distributed...  ...and document system design tradeoffs across consistency, performance, availability, and operational complexity Own backend services... 
    Senior
    Performance
    Full time

    Alterion, Inc.

    San Francisco, CA
    1 day ago
  •  ...500 companies and platform engineers across 100+ countries . Crossplane...  .... Upbound is hiring a Senior Software Engineer to help us build...  ...reconciliation problems, and performance bottlenecks. Communicate...  ...company embracing usage-based models. If you're excited to build... 
    Senior
    Performance
    Full time
    Remote work
    Worldwide

    Upbound

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...GraphQL. About the role Swayable is seeking a Senior Engineer blending Python software development expertise with scientific computing,...  ...improving our tools, techniques, and architecture for high-performance computing. You will work with a talented and diverse... 
    Senior
    Performance
    Full time

    Swayable

    San Francisco, CA
    1 day ago
  •  ...About the role This isn't just another engineering role. This is a unique opportunity to...  ...function and shape the future of performance and scalability at Persona. As our products...  ...What you'll bring to Persona A strong software engineering background, demonstrated by... 
    Senior
    Performance
    Full time
    Temporary work
    For contractors
    Internship

    Persona

    San Francisco, CA
    1 day ago
  •  ...Spring Health seeks a senior software engineer to enhance the Alma product’s search and matching capabilities. You will design scalable backend...  ...systems and microservices. This role focuses on performance, relevance, and maintainability at scale. #J-18808-Ljbffr
    Senior
    Performance

    Jobtailor

    San Francisco, CA
    8 hours ago
  • $195k - $225k

     ...MoM, and expanding the core engineering team in SF. The surface area...  ...enterprise The Role As the Senior Software Engineer – Backend (Systems...  ...deep ownership over APIs, performance, and data integrity — while...  ...backend feature (e.g., model queue, caching layer, or realtime... 
    Senior
    Performance
    Full time

    Vizcom

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...and collaborative group of software engineers, machine learning (ML) engineers...  ...measure and enhance the performance of the Waymo Driver. We...  ...achieve those goals by jointly modeling the real world, including...  ...role, you will report to a  Senior Staff Engineering Manager.... 
    Senior
    Performance
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $200k - $280k

     ...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud...  ...providing critical insights into system performance and GPU utilization, and proactively identifying...  ..., GPU clusters, and custom metrics for model performance and training pipelines.... 
    Senior
    Performance
    Full time
    Remote work

    Together Ai

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...improving the quality of the software that drives the car. We are...  ...experienced data-minded software engineers and data scientists to help...  ...signals to measure the performance and driving qualities of the...  ...data analysis tools for rapid modeling and prototyping... 
    Senior
    Performance
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...high-volume data replication simple, reliable, and scalable for engineering teams. Our platform powers mission-critical use cases...  ...and AI workloads. Artie is built for engineers who care about performance, reliability, and operational simplicity — and we’re growing... 
    Senior
    Performance
    Full time
    Visa sponsorship

    Artie

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...the infrastructure, tools, and model access that teams need to...  ...unified platform where high-performance inference, orchestration, and...  ...This role is ideal for engineers who thrive on complex distributed...  ...well-tested, and maintainable software and IaC for both new and existing... 
    Senior
    Performance
    Full time
    Currently hiring
    Relocation package

    Falò

    San Francisco, CA
    1 day ago
  •  ...Role Summary We have an opening for a Senior Software Engineer on our Infrastructure Team, with specific focus on Observability - both internal...  ...production readiness coverage for unit, integration, and performance of your feature ownership area. Own Set a high bar... 
    Senior
    Performance
    Full time

    Temporal Technologies

    San Francisco, CA
    1 day ago
  •  ...500k+ containers/month and growing. You’ll own reliability, performance, and security for multi‑tenant compute. What You’ll Do Design...  ...and enjoy tinkering with LLMs. Why Julius Small, senior team; massive impact surface; hard infra problems at meaningful... 
    Senior
    Performance
    Full time
    Remote work

    Julius Ai

    San Francisco, CA
    1 day ago
  • $196k - $220k

     ...playing games. We're looking for a data-driven full stack Senior Software Engineer to join the Growth team at Discord. Our team is...  ...high-impact opportunities and iterate quickly Maintain performance and reliability standards for user-facing web properties... 
    Senior
    Performance
    Full time

    Discord

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - Model Performance. Be the first to apply!