Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Global Inference Library Engineer

$175k - $250k

Australia-Employment

Global Inference Library Engineer $175000 - $250000 per year | San Francisco, CA | On-site | Permanent A bit about us: We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures. Why join us? Well-funded by leading tech investors Cutting edge technical problems with complex solutions Lucrative Equity in a seed stage startup Competitive compensation Excellent benefits (healthcare, vision, dental) Job Details We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments. What We’re Looking For Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software Strong programming experience with Python and C++, Rust, or similar systems languages Experience with LLM inference frameworks and model-serving infrastructure Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies Experience developing, integrating, or optimizing performance-critical compute kernels Understanding of modern transformer and LLM architectures Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management Experience benchmarking and profiling AI workloads across different hardware environments Strong understanding of GPU or accelerator architecture and performance characteristics Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable Jobot is an Equal Opportunity Employer. We provide an inclusive work environment that celebrates diversity and all qualified candidates receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, age (40 and over), disability, military status, genetic information or any other basis protected by applicable federal, state, or local laws. Jobot also prohibits harassment of applicants or employees based on any of these protected categories. It is Jobot’s policy to comply with all applicable federal, state and local laws respecting consideration of unemployment status in making hiring decisions. Sometimes Jobot is required to perform background checks with your authorization. Jobot will consider qualified candidates with criminal histories in a manner consistent with any applicable federal, state, or local law regarding criminal backgrounds, including but not limited to the Los Angeles Fair Chance Initiative for Hiring and the San Francisco Fair Chance Ordinance. Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Jobot, and/or its agents and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: jobot.com/privacy-policy #J-18808-Ljbffr Australia-Employment

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Global Inference Library Engineer in San Francisco, CA vacancy
  • $160k - $230k

     ...state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize...  ....We are seeking anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • $170k - $245k

     ...popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber,...  ...+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    5 days ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing... 
    Suggested

    Sail Research

    San Francisco, CA
    4 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate... 
    Suggested

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...A tech startup in AI model serving located in San Francisco is seeking a qualified candidate to architect scalable inference systems. The role focuses on optimizing model serving performance and integrating advanced techniques for AI deployment. Candidates should have... 
    Suggested

    Reducto, Inc.

    San Francisco, CA
    7 hours ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation... 
    Local area

    Inference

    San Francisco, CA
    3 days ago
  • $167.2k - $209k

     ...builders in the world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the...  ...batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    6 days ago
  •  ...Francisco, on‑site ABOUT THE ROLE You build and operate the inference systems that serve our models in production. The work spans serving...  ...that come with running real workloads. This is an engineering role, not a research role. You'll measure, profile, debug, and... 

    MakerMaker.AI

    San Francisco, CA
    6 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates... 
    Local area

    Mirai Labs

    San Francisco, CA
    3 days ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 

    inference.net

    San Francisco, CA
    4 days ago
  • Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference... 
    Remote work

    Inferact

    San Francisco, CA
    2 days ago
  •  ...Join a small, focused team of YC and unicorn founders and senior engineers with deep expertise in 3D, generative video, developer...  ...possible. About the Role We're looking for a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This... 
    Relocation
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    5 days ago
  • Help build the inference stack behind the next generation of multimodal foundation models. Most inference roles are about making existing...  ...production. You’ll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models... 
    Work at office
    Relocation package

    techire.®

    San Francisco, CA
    3 days ago
  • OpenArt is seeking a Growth Engineer focused on globalization for North America. You will own end-to-end globalization features, including multi-language infrastructure, region-specific onboarding, pricing experiments, and localization workflows. You will collaborate with... 

    Embedding VC

    San Francisco, CA
    5 days ago
  • OpenArt is seeking a Growth Engineer, Globalization to lead engineering projects that expand OpenArt's international reach. You’ll own end-to-end globalization features across localization, onboarding, pricing, checkout, and lifecycle communications. The role emphasizes... 

    Socket

    San Francisco, CA
    6 days ago
  • $300k - $400k

    Growth Engineer - Globalization (North America) About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools powered by cutting-edge AI, enabling anyone to create videos, visuals... 
    Worldwide
    Visa sponsorship

    Socket

    San Francisco, CA
    6 days ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 

    Anyscale

    San Francisco, CA
    4 days ago
  •  ...Its APA system, built on the industry’s first Process Reasoning Engine (PRE) and specialized AI agents, combines process discovery, RPA...  ...intersection of technical excellence and executive strategy, owning the global pre-sales function and serving as a key voice in shaping company... 
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Aisera

    San Francisco, CA
    4 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates... 
    Full time

    Vast.ai

    San Francisco, CA
    2 days ago
  •  ...infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with #J-18808-... 

    Magic AI, Inc

    San Francisco, CA
    4 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    2 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  • $293k - $385k

     ...seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by...  ...AI/ML workloads, including training or inference systems.Familiarity with GPU or accelerator...  ...requests can be made via this link.OpenAI Global Applicant Privacy PolicyAt OpenAI, we... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm... 

    Kindredventures

    San Francisco, CA
    4 days ago
  • Lunar Energy in San Francisco, CA is seeking an IT Support Specialist to join a fast-growing team supporting a global workforce. You will provide 1st/2nd line device and SaaS support, handle onboarding/offboarding, procure equipment, and optimize IT processes while leveraging... 

    Lunar Energy

    San Francisco, CA
    4 days ago
  • $225k

    Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate... 

    Dormont Manufacturing Co

    San Francisco, CA
    3 days ago
  • $112.5k - $147.5k

    DescriptionKforce has a client that is seeking an Electric Global Supply Manager in San Francisco, CA.About the Role:The client is...  ...development and manufacturing. This role will collaborate closely with engineering and manufacturing to execute our electrical roadmap.Duties:*... 
    Flexible hours

    KForce

    San Francisco, CA
    17 hours ago
  • $420k

     ...real-world industry insight, drawing on decades of experience in engineering, construction management, scheduling and delay analysis, and...  ...during project execution.This is an opportunity to work on globally significant matters within a team known for its strategic thinking... 
    For contractors
    Worldwide

    Jameson Legal

    San Francisco, CA
    4 days ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and optimizations to ensure high throughput and low latency in production. You will collaborate... 

    MakerMaker

    San Francisco, CA
    2 days ago
  •  ...is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving...  ...components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly... 

    Gimlet Labs

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Global Inference Library Engineer. Be the first to apply!