Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Global Inference Library Engineer

$175k - $250k

Jobot

This Jobot Job is hosted by: Grant Greenhalgh Are you a fit? Easy Apply now by clicking the "Quick Apply" button and sending us your resume. Salary: $175,000 - $250,000 per year A bit about us: We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures. Why join us? Well-funded by leading tech investors Cutting edge technical problems with complex solutions Lucrative Equity in a seed stage startup Competitive compensation Excellent benefits (healthcare, vision, dental) Job Details We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments. What We’re Looking For Strong experience building AI/ML infrastructure, inference systems, or high-performance computing software Strong programming experience with Python and C++, Rust, or similar systems languages Experience with LLM inference frameworks and model-serving infrastructure Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies Experience developing, integrating, or optimizing performance-critical compute kernels Understanding of modern transformer and LLM architectures Familiarity with inference concepts including batching, attention, KV caching, quantization, and memory management Experience benchmarking and profiling AI workloads across different hardware environments Strong understanding of GPU or accelerator architecture and performance characteristics Experience with frameworks such as vLLM, TensorRT-LLM, SGLang, or similar inference technologies is highly valuable Interested in hearing more? Easy Apply now by clicking the "Quick Apply" button. Jobot is an Equal Opportunity Employer. We provide an inclusive work environment that celebrates diversity and all qualified candidates receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, age (40 and over), disability, military status, genetic information or any other basis protected by applicable federal, state, or local laws. Jobot also prohibits harassment of applicants or employees based on any of these protected categories. It is Jobot’s policy to comply with all applicable federal, state and local laws respecting consideration of unemployment status in making hiring decisions. Sometimes Jobot is required to perform background checks with your authorization. Jobot will consider qualified candidates with criminal histories in a manner consistent with any applicable federal, state, or local law regarding criminal backgrounds, including but not limited to the Los Angeles Fair Chance Initiative for Hiring and the San Francisco Fair Chance Ordinance. Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Jobot, and/or its agents and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: jobot.com/privacy-policy

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Global Inference Library Engineer in San Francisco, CA vacancy
  • $230k - $385k

     ...team combines world-class researchers, engineers, designers, and operators who care deeply...  ...Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the...  ...can be made via this link . OpenAI Global Applicant Privacy Policy At OpenAI,... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    3 days ago
  • $170k - $245k

     ...popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber,...  ...+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  •  ...through multiple rounds of revisionKeep files, versions, and asset libraries organized and correctly namedCollaborate with communications,...  ...of eligible wages and dependent on company performance.Global benefits provide options for the following:Paid Time Off: earned... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Freelance
    Internship
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    San Francisco, CA
    14 hours ago
  • $119.88k - $159.84k

     .... Anticipated Posting Close Date*: 12/22 *Job posting may close early due to the volume of applicants. Global AV Events and Production Engineer Fastly is a leading technology company dedicated to providing a better, faster internet experience. We're seeking... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Fastly

    San Francisco, CA
    2 days ago
  • $347k

     ...seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by...  ...AI/ML workloads, including training or inference systems.Familiarity with GPU or accelerator...  ...requests can be made via this link.OpenAI Global Applicant Privacy PolicyAt OpenAI, we... 
    Suggested
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $420k

     ...real-world industry insight, drawing on decades of experience in engineering, construction management, scheduling and delay analysis, and...  ...during project execution.This is an opportunity to work on globally significant matters within a team known for its strategic thinking... 
    For contractors
    Worldwide

    Jameson Legal

    San Francisco, CA
    3 days ago
  • $182k - $242k

     ...Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior...  ...: We're looking for a Senior Storage Engineer, File & Block to help build and operate...  ...that our largest AI training and inference workloads and internal stateful services... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    3 days ago
  • $235k - $260k

     ...Open Source Engineer Firecrawl is looking for a high-agency engineer to own our open source work. That starts with the main Firecrawl...  ...superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now. What You'll Do... 
    Full time
    Temporary work
    For contractors
    Remote work
    Work from home
    Visa sponsorship
    Flexible hours

    Firecrawl

    San Francisco, CA
    2 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    3 days ago
  • $90k - $105k

     ...0.00/yr - $105,000.00/yr Implementation Engineer / Full Stack Engineer This is a hybrid (...  ...integrated business apps. Odoo has become a global network with more than 12+ million users...  ...chances of interviewing at Odoo by 2x Inferred from the description for this job... 
    Full time
    Immediate start
    Remote work
    Work from home
    Home office
    Work visa

    Odoo

    San Francisco, CA
    3 days ago
  • About The Role As a Growth Engineer, you will own how aion attracts, engages, and activates...  .... Headquartered in the U.K., we operate globally with core teams in London, SF, and NYC....  ...: comfortable discussing training vs. inference with developers Social‑native (AI Twitter... 
    Work at office
    Remote work
    Flexible hours
    Night shift

    aion

    San Francisco, CA
    6 days ago
  • $300k

    GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real...  ...The company is revenue‑generating, its models are used by global enterprises, and the SF R&D team is expanding following a... 
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    5 days ago
  •  ...future of enterprise AI. We’re hiring a GTM Engineer to help design, operate, and...  ...of AI infrastructure, from low-latency inference to scalable model serving. Build What’s...  ...how businesses and developers harness AI globally. Ownership & Impact: Join a fast-growing... 

    Fireworks AI

    San Francisco, CA
    5 days ago
  • $170k - $230k

     ...re on a mission to build the engine of superintelligence, and developers...  .... This role will support the global developer relations strategy,...  ...fine-tuning, and production inference), and familiarity with modern...  ...open-source frameworks, libraries, and agent-based systems.... 
    Full time
    Flexible hours

    Nscale

    San Francisco, CA
    3 days ago
  • $187.5k - $247.5k

     ...Redwood Materials Redwood is localizing a global battery supply chain that seamlessly...  ...have. Staff Mechanical Design Engineer, EPC Redwood Materials is hiring...  ...professional or employment information, and inferences drawn from your PI. We collect your PI... 
    Full time
    Work experience placement

    Redwood Materials

    San Francisco, CA
    3 days ago
  • $153k - $248.5k

     ...Redwood Materials Redwood is localizing a global battery supply chain that seamlessly...  ...batteries we already have. Functional Safety Engineer, Energy Storage Redwood Materials is...  ...professional or employment information, and inferences drawn from your PI. We collect your PI... 
    Full time
    Casual work
    Work at office
    Local area
    Remote work
    Night shift

    Redwood Materials

    San Francisco, CA
    14 hours ago
  • $190k - $240k

     ...At Ouster, we build sensors and tools for engineers, roboticists, and researchers, so they can make the world safer and more efficient....  ...the company and need your help! We are looking for a Director, Global Field Application Engineer to build and lead the team that sits... 
    Full time
    Work experience placement
    Local area
    Worldwide

    Ouster

    San Francisco, CA
    1 day ago
  • $200k - $250k

     ...to find a high-performing Vision Sales Engineer. If you want to be at the forefront of the...  ...cameras on production lines and running inference on the edge. This is not a startup...  ...repeatable rollouts across lines, plants, and global sites. Who You Are Hands-On... 
    Full time

    TRC Talent Solutions

    San Francisco, CA
    21 days ago
  • $152k - $200k

     ...Redwood Materials Redwood is localizing a global battery supply chain that seamlessly...  ...for a Battery Software Integration Engineer to build the bridge between diverse, second...  ...professional or employment information, and inferences drawn from your PI. We collect your PI... 
    Full time

    Redwood Materials

    San Francisco, CA
    3 days ago
  • $157k - $302k

     ...generation of frontier model training and inference workloads.The Hardware Operations team...  ...Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure,...  ...expanding AI environments.As we scale globally, we are building the operational frameworks... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  •  ...looking for a computer vision & deep learning engineer to advance the state of the art in low-...  ..., train models, and deploy real-time inference on embedded hardware. You’ll utilize...  ...Understanding of CMOS imaging (rolling vs. global shutter, read noise, dynamic range,... 
    Contract work
    Work experience placement

    Socket

    San Francisco, CA
    2 days ago
  • $342k

     ...specifically for AI.About the RoleAs an Engineer on our hardware optimization and co-design...  ...towards efficient training and inference on our models. If you are excited about...  ...requests can be made via this link.OpenAI Global Applicant Privacy PolicyAt OpenAI, we believe... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $125k - $170k

     ...Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and...  ...cleanup and improvement of internal systems across our growing global team. This is not a traditional help desk role, and it is... 
    Full time
    Work at office
    Remote work

    Inferact

    San Francisco, CA
    10 days ago
  • $160k - $252.5k

     ...Redwood Materials Redwood is localizing a global battery supply chain that seamlessly...  ...the past. Power Electronics Controls Engineer As a Power Electronics Controls...  ...professional or employment information, and inferences drawn from your PI. We collect your PI... 
    Full time
    Shift work

    Redwood Materials

    San Francisco, CA
    4 days ago
  • $170k - $205k

     ...are seeking a highly skilled, hands-on Senior level performance engineer who thrives on deep technical challenges. The ideal candidate...  ...benchmarking and system tuning, making a tangible impact on Crusoe’s global compute operations.What You'll Be Working On:System Performance... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    14 hours ago
  • $134.5k - $265.1k

    Position Summary Senior Consultant - Forward Deployed Engineer (FDE) - DatabricksForward Deployed Engineers at Deloitte work alongside...  ...in the recruiting process, please direct your inquiries to the Global Call Center (GCC) at ****@*****.***.... 
    Local area

    Deloitte

    San Francisco, CA
    4 days ago
  • $115k - $160k

     ...Technology and Workflow EngineerCooley is seeking a Collaboration Technology and Workflow Engineer to join the Technology Platforms and Applications team.About Cooley: Cooley is a global law firm with an expansive practice and more than 3,000 employees and partners... 
    Full time
    Temporary work
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours
    Weekend work

    Cooley

    San Francisco, CA
    2 days ago
  • $100k - $120k

     ...teams collaborate, present, and make decisions every day.As our AV Engineer, you’ll own the technology that powers those moments: designing...  ...AV lead for new site expansions as Crusoe continues to grow globally. We’re looking for someone who thinks like a mountaineer about... 
    Temporary work
    For contractors
    Work at office
    Remote work
    Night shift

    Crusoe

    San Francisco, CA
    3 days ago
  • $220k - $320k

     ...shape a brighter way forward.Role OverviewAs a Forward Deployed Engineer, you will operate at the front lines of innovation — embedded...  ...rooted in Infrastructure initiatives reporting directly to the Global Head of Infrastructure Development.Key ResponsibilitiesSolution... 
    Full time
    Local area
    Remote work
    Shift work

    Jones Lang LaSalle

    San Francisco, CA
    1 day ago
  • $140k - $180k

     ...Marketing /Full-time /On-siteAbout FinixFinix is building the global operating system for fintech, starting with payments. From startups...  ...the RoleWe’re looking for our first Go-To-Market (GTM) Engineer to architect and operationalize the systems behind Finix’s pipeline... 
    Full time
    Visa sponsorship
    Free visa
    Flexible hours

    Finix Payments

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Global Inference Library Engineer. Be the first to apply!