Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Model Serving & High-Performance API Engineer

Black Forest Labs Inc.

Black Forest Labs Inc. is seeking a Member of Technical Staff to bridge research breakthroughs and production systems. You will own the bridge between research checkpoints and production-ready inference services, designing scalable APIs and high-performance serving. You will optimize GPU workloads, manage distributed systems, and collaborate with researchers to ship demos and live endpoints rapidly, across cloud infrastructures. #J-18808-Ljbffr Black Forest Labs Inc.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Model Serving & High-Performance API Engineer in San Francisco, CA vacancy
  •  ...bring cutting-edge models into production...  ...the platform engineers turn to to ship...  ...’s Model Performance (MP) team is responsible...  ...on Model API's — the infrastructure...  ...systems, model serving, and developer...  ...join a small, high‑impact team...  ...and curiosity. ML experience is a... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $180k - $300k

     ...’re creating the generative models that power how people make images...  ...before becoming usable APIs Inference is slower than it...  ...services Design and maintain high-performance APIs serving millions of requests...  ...performance, and production ML serving. What We’re Looking... 
    Performance
    Remote work
    Worldwide
    2 days per week

    Black Forest Labs

    San Francisco, CA
    1 day ago
  •  ...experienced backend/infrastructure engineer to build platforms powering AI workloads, including model training, serving, and vector search. You will join a high-visibility team connected to research...  ...across platform, infra, and ML teams to deliver end-to-end experiences... 
    Suggested

    Databricks

    San Francisco, CA
    1 day ago
  •  ...leader in foundational AI models for image and video. Based...  ...in San Francisco, we seek engineers who can bridge research breakthroughs...  ...You will design scalable APIs and optimize GPU inference...  ...backend systems, GPU performance, and production ML serving, offering a hybrid setup... 
    Performance

    Black Forest Labs

    San Francisco, CA
    4 days ago
  • $166k - $225k

     ...business. Databricks’ Model Serving product provides...  ...and manage AI/ML models — from traditional...  ....As a Senior Engineer, you’ll play a critical...  ...that enable high-throughput, low-latency...  ...core systems and APIs that power...  ...-offs to optimize performance, throughput, autoscaling... 
    Performance
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $192k - $260k

     ...their business. Foundation Model Serving is the API Product for hosting and...  ...For this role, no prior ML or AI experience is necessary. We’re looking for engineers who have owned high scale operational sensitive...  ...trade-offs to optimize performance, throughput, autoscaling,... 
    Performance
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $190k - $265k

     ...business. Founded by engineers — and customer-...  ...apps, AI agents, model training, model serving, and Vector Search...  ...You'll be joining a high-agency, high-...  ...grade reliability and performance. Our Foundation Model APIs provide a unified...  ...platform, infra, and ML teams to deliver... 
    Performance
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $325k

     ...company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments....  ...experience, strong familiarity with ML architectures, and experience...  ...collaboration with researchers and focus on performance optimization. Compensation ranges... 
    Performance

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...We’re a team of engineers, clinicians, and...  ...helps care teams perform with greater...  ...and integrating high performance robotic...  ...research, SW/ HW/ ML engineering,...  ...Responsibilities• GPU & Model Performance &...  ...inference—while serving as a core...  ...in GPU Compute API - CUDA, OpenCL•... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • $166.6k - $208.3k

     ...scope and in stakes. Models increasingly...  ...owns the production ML lifecycle: the...  ...themselves — and we serve low-latency, highly available scores to the decision engine that depends on...  ...versioning, CI/CD with performance, bias, and...  ...in Python, with API frameworks like FastAPI... 
    Performance

    Mercury

    San Francisco, CA
    3 days ago
  •  ...AI, large language models, and our...  ...looking for a Founding Engineer & CTO to be the technical...  ...production AI/ML systems, including...  ...team, fostering a high-performance, collaborative culture...  ...third-party APIs, data providers, and...  ...tenant SaaS platforms serving enterprise clients... 
    Performance
    Flexible hours

    NovusMinds AI

    San Francisco, CA
    1 day ago
  • $220k - $320k

     ...every last drop of performance out of GPUs,...  ...specialized language models for companies that...  ...ten‑person team of engineers who work in‑person...  ...Francisco on difficult, high‑impact engineering...  ...with the goal of serving models faster and...  ...with applied ML engineers to ensure... 
    Performance
    Work at office

    SOLANA FOUNDATION

    San Francisco, CA
    22 hours ago
  • $204k - $348k

     ...Principal/ Principal Software Engineer, AI Lab Execution...  ...this role, you will serve as a technical leader...  ...interfaces, services, high-performance APIs, databases, and reliability...  ...'ll work closely with ML researchers, platform...  ...Design systems that model scientific intent,... 
    Performance
    Full time
    Work at office
    Local area
    Flexible hours

    Lila Sciences

    San Francisco, CA
    4 days ago
  • $180k - $225k

    As a Software Engineer on the ML Infrastructure team, you will...  ..., and efficient serving of LLMs. Our platform...  ...design. You’ll work in a highly collaborative environment...  ...fault-tolerant, high-performance systems for serving...  ...and optimize models for production and research... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $86.7k - $170.9k

     ...Summary Integration Engineer II, AI & Engineering/...  ...improve financial performance, accelerate new digital...  ...platforms. Our delivery models are tailored to meet each...  ...relationships. You will be part of API management strategies...  ...industries/sectors you serve. Must be legally... 
    Performance
    Local area
    Flexible hours

    Deloitte

    San Francisco, CA
    1 day ago
  •  ...talented Robotics Software Engineer to design and build...  ...to physical robots performing customer‑critical tasks...  ...control interfaces, and high‑level APIs that power real‑world...  ...Integrate and deploy ML policies (PyTorch) into...  ...distributed training/serving infrastructure. Take... 
    Performance
    Full time
    Work experience placement
    Immediate start

    Verne Robotics

    San Francisco, CA
    5 days ago
  •  ...AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal candidate has over... 
    Performance

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $300 per month

     ...strategies, and be part of a high-performing team that believes in...  ..., our Production Engineering team ensures the...  ...experience with large language models to help us build and...  ...with a focus on serving and scaling LLM workloadsBuild...  ...models (LLMs) or AI/ML infrastructureSRE... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency...  ...to craft end-to-end serving pipelines and push the boundaries... 
    Performance

    Together AI

    San Francisco, CA
    2 days ago
  • $218.4k - $273k

     ...Robotics and developing ML pipelines for training...  ...of Robotics data and model evaluation. We are...  ...Senior Machine Learning Engineer, Computer Vision to...  ...(MOT/SOT): Designing high-performance deep learning models for...  ...Technical Leadership: Serve as the subject matter... 
    Performance
    Full time

    Scale Ai

    San Francisco, CA
    1 day ago
  • $225k - $275k

    ML Engineer - AI-Powered Automation & Workflow Intelligence High‑velocity tech startup at the intersection...  ...ML systems that serve as the...  ...involving large language models (LLMs),...  ...reproducibility and performance standards while integrating...  ...scalable APIs or backend services... 
    Performance
    Full time

    Blue Signal Search

    San Francisco, CA
    2 days ago
  •  ...unicorn founders and senior engineers with deep expertise in 3D,...  ...for a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is a...  ...performance from generative media models. You'll work across the model-serving stack, designing novel... 
    Performance
    Relocation
    Visa sponsorship
    Relocation package

    Reactor.am

    San Francisco, CA
    5 days ago
  • $350k

     ...We are scientists, engineers, and builders who’ve...  ....ai, open-weights models like Mistral, as well...  ...'ll join a small, high-impact team...  ...research, and ultimately serve AI models and build...  ...controllers/operators, or performance profiling. Familiarity with GPU/ML workflows or large‑... 
    Performance
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Thinking Machines Lab

    San Francisco, CA
    1 day ago
  • $170k - $200k

     ...training robotics foundation models. Instawork Robotics is...  ...Instawork Robotics ML Engineer will help build and...  ...the quality and performance of our dataset and ML...  ...for ActionWe practice high-velocity decision-making...  ...Our AI-powered platform serves thousands of businesses... 
    Performance
    Hourly pay
    Internship
    Local area
    Shift work

    Instawork

    San Francisco, CA
    4 days ago
  •  ...bring cutting-edge models into production....  ...build the platform engineers turn to to ship AI...  ...for Disaggregated Serving, Wide Expert Parallelism...  ...networking performance on bleeding-edge clusters...  ...experience with high-performance...  ...Exposure to a variety of ML startups, offering... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $248.4k - $310.5k

     ...Software Engineer - Robotics & Autonomous Systems...  ...data collection, model training pipelines...  ...datasets Build ML training and fine-...  ...features at high velocity while maintaining...  ...reliability and performance Ideally, You...  ...model deployment and serving frameworks... 
    Performance
    Full time

    Scale Ai

    San Francisco, CA
    1 day ago
  • $98k - $140k

     ...You’ll work with product and engineering teams to build systems to define...  ...to deliver reliable and high-quality AI experiences. Your work...  ...of that you'll shape Notion's model strategy and work directly with...  ...working with data — You can self-serve insights from large datasets,... 
    Live in
    Local area

    Notion Labs

    San Francisco, CA
    2 days ago
  • $144k - $164k

     ...Management, Gen AI Model Gateway At...  ...development, and serving). The FM Gateway...  ...requirements and deliver high value solutions...  .../or shipping AI/ML-powered products...  ..., or software engineering. Preferred...  ...experience with API or AI Gateway services...  ...hired to perform work within one... 
    Full time
    Part time
    Local area

    Capital One National Association

    San Francisco, CA
    4 days ago
  •  ...and Amsterdam. Sales Engineering unlocks and empowers prospects...  .... Responsibilities Serve as the technical lead...  ..., and rectify any performance shortfalls Constantly...  ...statistical/ML models, translating analytical...  ...specific to interfacing with API platforms Willing and... 
    Performance
    Work experience placement
    Local area

    Plaid Inc

    San Francisco, CA
    3 days ago
  •  ...are forming small, highly capable product...  ...management, design, and engineering together with the...  ..., data pipelines, APIs, identity...  ...integration failures, model behavior problems,...  ...data quality issues, performance bottlenecks, and...  ...approximately 791,000 people serving clients in more... 
    Performance
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Model Serving & High-Performance API Engineer. Be the first to apply!