Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI/ML Infra / Systems Engineer (Inference)

$204k - $216k

Sapience AI Corporation

Job Description

Job Description

About Sapience AI

Sapience AI is the collective intelligence platform for professional communities. We sit above the CRMs, AMS platforms, and knowledge bases that organizations already run, and we turn the expertise scattered across them into something every member can search, act on, and share.

The intelligence a community needs is already inside it. Most organizations just cannot reach it. Knowledge lives in silos, in legacy systems, in the heads of a few experts, and in fragmented records no one can connect. We change that.

Our work is grounded in four commitments: technology elevates people and never replaces them, the best expertise is already inside the community, everything is built on trust, and every deployment is purpose-driven for the organization it serves.

Let's achieve more, together.

Where this role sits

This role owns how Sapience AI serves intelligence at scale. You build and run the inference infrastructure that turns the models and reasoning behind MINERVA and COGENT into fast, reliable, affordable answers for members.

You work where models meet production: serving, scaling, latency, and cost. Every time a member asks the platform something, your systems are what make the answer arrive quickly and hold up under load.

You sit close to applied AI, research, and platform engineering, and you make the difference between a model that works in a notebook and one that serves a community reliably.

Why this role exists

Collective intelligence is only useful if it is fast and dependable. Members will not wait, and communities cannot rely on a platform that buckles under load or costs too much to run.

Serving modern AI at scale is a hard systems problem: large models, tight latency budgets, expensive hardware, and demand that spikes. It takes real infrastructure engineering to get right.

The AI/ML Infrastructure and Systems Engineer owns that problem. You make inference fast, reliable, and affordable, so the intelligence the platform promises actually reaches members.

What you will own (Areas of Responsibility)

You hold seven areas of responsibility across inference infrastructure. Each one is yours to set direction on, build, and measure.

1. Inference serving and systems
  • Build and operate the inference systems that serve models and reasoning behind MINERVA and COGENT.
  • Own the serving path end to end, from request to response, under real load.
  • Make serving robust to failure, traffic spikes, and change.
2. Latency and performance
  • Drive down latency so members get fast answers, and keep it low as the platform grows.
  • Optimize the full path, including model execution, batching, caching, and retrieval.
  • Profile relentlessly and remove the bottlenecks that matter.
3. Scale and reliability
  • Scale inference to more members, more communities, and heavier reasoning without losing reliability.
  • Build autoscaling, load management, and graceful degradation.
  • Own the reliability of the serving layer as a first-order responsibility.
4. Cost and efficiency
  • Own the economics of inference, including GPU and hardware efficiency at scale.
  • Improve utilization and cut waste without hurting quality or speed.
  • Make cost a designed property, not a surprise.
5. Model deployment and lifecycle
  • Build the paths that take models and reasoning components from research to production safely.
  • Support rollout, versioning, and rollback with confidence.
  • Give applied AI and research a fast, safe route to ship.
6. Observability and operations
  • Instrument the serving layer so it can be measured, debugged, and trusted.
  • Build the alerts, dashboards, and tooling that keep inference healthy.
  • Reduce the operational burden of running AI at scale.
7. Hardware, accelerators, and platform choices
  • Make sound choices about accelerators, runtimes, and serving frameworks.
  • Balance performance, cost, and maintainability in platform decisions.
  • Keep the stack current as inference technology moves.
AI-augmented ways of working

You build the infrastructure that serves AI, and you use AI in building it, to generate tooling, reason about performance, and move faster, while you own correctness, reliability, and cost.

The standard is human in partnership: AI accelerates the work, you own the judgment, the interpretation, and the call. The people who create the most value here are not the ones producing the most output. They are the ones turning evidence into clear, durable decisions.

What this role is not

To keep the boundary clear:

  • This is not a model research role. You serve and scale models; you do not develop new model architectures.
  • This is not a general backend role. Your center of gravity is inference systems, performance, and hardware efficiency.
  • This is not a data engineering role. You partner with data teams, but you own serving, not the data platform.
  • This is not a best-effort prototype role. You are accountable for production inference members depend on.
What success looks like

We measure this role on outcomes the team can see:

  • Fast answers. Members get low-latency responses, and latency stays low as the platform grows.
  • Reliable at scale. Inference holds up under load, spikes, and growth.
  • Affordable intelligence. Inference cost per answer improves as usage rises.
  • Safe rollouts. Models and reasoning components ship, version, and roll back with confidence.
  • Operable serving. The serving layer is instrumented, debuggable, and healthy.
  • Sound platform choices. Accelerator and framework decisions age well.
Who you are Required qualifications
  • Five or more years in infrastructure, systems, or ML infrastructure engineering.
  • Hands-on experience serving ML or LLM models in production at scale.
  • Deep understanding of latency, throughput, and performance optimization.
  • Experience with GPUs or accelerators and their efficient use.
  • Strong systems programming and distributed systems fundamentals.
  • A track record of reliable, cost-aware production systems.
  • Fluency with observability and operational excellence.
Preferred qualifications
  • Experience with inference-serving frameworks and model runtimes.
  • Experience optimizing LLM inference, including batching, quantization, and caching.
  • Familiarity with retrieval systems and their performance characteristics.
  • Experience owning cost and capacity for AI workloads.
  • Exposure to neuro-symbolic or agentic systems in production.
How you work
  • You name the real bottleneck before reaching for a fix.
  • You measure before and after, and you trust evidence over intuition.
  • You treat reliability and cost as first-order, alongside speed.
  • You build systems others can operate.
  • You share tooling and knowledge across the team.
Skills & Competencies
  • Inference serving architecture and optimization.
  • Latency, throughput, and performance engineering.
  • GPU and accelerator efficiency.
  • Distributed systems, scaling, and reliability engineering.
  • Model deployment, versioning, and rollout.
  • Observability and production operations for AI.
  • Cost and capacity management for inference.
Services & Tools Experience
  • Inference-serving frameworks and model runtimes (for example vLLM, TensorRT, Triton-class systems).
  • GPU tooling, CUDA-class ecosystems, and accelerator runtimes.
  • Kubernetes, containers, and cloud platforms (AWS, GCP, or Azure).
  • Observability stacks (metrics, tracing, and logging).
  • Python and a systems language such as Go, Rust, or C++.
  • Caching, queuing, and load-management systems.
  • Serving the models and reasoning behind the MINERVA platform and COGENT architecture.
Prior Experience & Background
  • Prior ML infrastructure, platform, or systems engineering at a software or AI company.
  • Experience serving models in production under real load.
  • A track record of performance and reliability improvements at scale.
  • Experience owning inference cost is a plus.
Cross-functional partners

You work most closely with Applied AI, Neuro-Symbolic AI, Platform Engineering, and Research. You serve the models and reasoning that power MINERVA and the COGENT architecture, and you own how they run in production.

How we hire

We review every application, and we encourage you to apply even if you do not match every line above. Research shows that talented people, especially those from underrepresented communities, often hold back when they do not meet every qualification. If that is the only thing holding you back, apply anyway.

Sapience AI is an equal opportunity employer. We are committed to a workplace where everyone, regardless of background, has a voice in building what comes next.

Compensation

Base Salary:  $204,000 - $216,000 + early stage equity

Generous health and wellness benefits

 

Sapience AI is an equal opportunity employer. We do not discriminate on the basis of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. If you need an accommodation to complete our application process, let your recruiter know.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI/ML Infra / Systems Engineer (Inference) in San Francisco, CA vacancy
  • $350k

     ...AI Systems Engineer Salary range: $350,000 - $600,000/year + benefits...  ...and development of our core ML stack, building systems that...  ...Interpretability: Inference stacks that are as performant...  ...building and path-set on what infra we should build Help other... 
    Suggested
    Visa sponsorship
    Flexible hours

    Transluce

    San Francisco, CA
    4 days ago
  •  ...investing deeply in Generative AI — pioneering advanced...  ...Machine Learning Systems Engineer (P60) to lead technical...  ...ll Doð Design and Build ML SystemsArchitect and...  ...data.Develop tools and infra to support rapid experimentation...  ...-scale model training, inference pipelines, or search/... 
    Suggested
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  •  ...is building the design suite for molecules, training frontier AI/ML models to advance drug discovery. You will develop core frameworks...  ...training and evaluation, in collaboration with researchers and engineers. Your role focuses on creating a scalable, reliable training... 
    Suggested

    Jobleads-US

    San Francisco, CA
    22 hours ago
  • Senior AI Engineer, MLOps & Distributed SystemsThe OpportunityJoin the...  ...Notifications pod is creating the systems that determine what message a...  ....You’ll work closely with ML Scientists, Data Scientists,...  ...reliable online and batch inference capabilities that meet clear... 
    Suggested
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Worldwide
    3 days per week

    Hinge Health

    San Francisco, CA
    1 day ago
  •  ...Quantum superintelligence is an AI that uses quantum computers...  ...and superconducting‑qubit systems, turning raw cryogenic hardware...  ...most of the world's software engineers. AI is already generating quantum...  ...collection, labelling, and inference. Integrate with external... 
    Suggested

    Conductor Quantum

    San Francisco, CA
    1 day ago
  •  ...other priorities. We can hire people in any country where we have a legal entity. Responsibilities As a ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model serving systems end-to-end. You will have the chance... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  •  ...reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for...  ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...systems. About the role: The Safeguards ML Infra team designs, builds, and operates the... 
    Full time

    Anthropic

    San Francisco, CA
    23 days ago
  • Join a leading AI lab's cutting-edge GenAI team to be at the core...  ...We're seeking talented MLOps Engineers with deep, hands-on expertise...  ...training data for frontier AI systems. This is a W-2 employment position...  ...training infrastructure, and ML framework-level topics. Design... 
    Full time
    Weekday work

    Obsidian

    San Francisco, CA
    4 days ago
  •  ...Job Description Full Stack Engineer, AI Systems, Artificial Intelligence (AI) Required, Work From Home   We are looking for a Full...  ...observability, and fallback mechanisms. - Collaborate closely with ML, backend, and product teams to ship features end-to-end. -... 
    Full time
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    10 days ago
  • $250k - $280k

     ...Description Staff / Principal Founding Engineer (Backend-Leaning) – AI Systems Platform San Francisco (in-office...  ..., TypeScript, APIs, AWS, cloud infra) Background in 0→1 or early-stage...  ...systems, agent frameworks, or data/ML pipelines, come from top 5 CS undergrad... 
    Work at office
    Immediate start
    Flexible hours

    Xpertalent

    San Francisco, CA
    a month ago
  •  ...Capital One in San Francisco is seeking a Staff AI Engineer to build responsible AI systems, partnering with engineers, researchers, technical program...  ...Capital One. The role covers foundation model training, inference, guardrails, evaluation, and observability, using Open... 

    Capital One

    San Francisco, CA
    1 day ago
  • $124k - $280k

    At PwC, our people in data and analytics engineering focus on leveraging advanced...  ...on developing and implementing advanced AI and ML solutions to drive innovation and enhance...  ...and optimising algorithms, models, and systems to enable intelligent decision-making and... 
    H1b

    PwC

    San Francisco, CA
    4 days ago
  • $300k

     ...startup building out their AI and cloud platform, powered...  ...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability...  ...frontier AI workloads, automate systems at petascale, and be part...  .... Collaborate with ML, networking, and platform teams... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $300k - $400k

     ...Mirendil Inc. is seeking a Systems Engineer in San Francisco to locate bottlenecks and optimize how compute is utilized. You will influence AI workload design, system co-design, and data movement across machines to maximize goodput while ensuring correctness under load... 

    Jobleads-US

    San Francisco, CA
    22 hours ago
  •  ...Technical Staff with expertise in generative modelling to work in a collaborative, hybrid setting. You’ll develop and deploy advanced AI models for biotech and pharmaceutical applications. Qualified candidates will have significant experience in machine learning and be... 

    Latent Labs

    San Francisco, CA
    2 days ago
  • $110 per hour

     ...technical talent with leading AI research labs. Headquartered in...  ...Dorsey . Position: MLOps Engineer (JAX, PyTorch, Pallas/Triton)...  ...training infrastructure, and ML framework-level topics . Design...  ...solutions to MLOps and ML systems problems . Evaluate MLOps... 
    Contract work
    Summer work
    Remote work
    Weekday work

    Mercor

    San Francisco, CA
    3 days ago
  • $140k - $210k

     ...destiny. Klaviyo is building an AI-first company-and that starts...  ...for a Lead People Technology Engineer to be the first dedicated AI...  ...design and build intelligent systems that transform how we hire,...  ...experience, with 2+ years building AI/ML or LLM-powered delivering... 
    Shift work

    Klaviyo

    San Francisco, CA
    3 days ago
  •  ...interpretable, and steerable AI systems. We want AI to be safe and beneficial...  ...of committed researchers, engineers, policy experts, and business...  ...come from understanding the ML workload well enough to know...  ...running ML training or inference infrastructure at scale... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    4 days ago
  •  ...purpose and we are hiring the world’s best engineers, scientists, designers, product managers,...  ...the whole company, and you decide how AI runs here rather than inheriting someone...  ...that connect these platforms to internal systems, with least-privilege access and full audit... 

    Juul

    San Francisco, CA
    2 days ago
  • $250k

     ...Join a rapidly scaling AI cloud infrastructure provider...  ..., experimentation, and inference at scale. The company...  ...Staff Site Reliability Engineer to support and scale...  ...closely with platform, ML, and infrastructure...  ...available infrastructure systems Improve CI/CD pipelines... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...of day 2 operations for our ML infrastructure platform. You...  ...envision and build robust systems, processes, automations, and... 
    Flexible hours

    Baseten

    San Francisco, CA
    4 days ago
  •  ...A leading AI research firm in San Francisco is seeking experienced signal integrity system design engineers. Responsibilities include leading system signal integrity design for AI supercomputing products and collaborating with multi-disciplinary engineering teams. Ideal... 
    Relocation package

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...Senior Systems Engineer San Francisco, California Onsite or Remote At Evidently, we are raising the quality of care for every patient...  ..., security and compliance posture, database performance, AI inference infrastructure: you can cover these areas when they touch systems... 
    Remote work
    Work from home

    Evidently

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems that can interact with...  ...abstractions Build update rules, inference mechanisms, and dynamic graph...  ...ensure symbolic layers work alongside ML, RL, and systems architecture... 
    Full time
    Immediate start

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  •  ...About the Team The Recursive Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build AI systems that accelerate and ultimately conduct high-quality research at OpenAI. We work to automate real research workflows and improve... 

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...Product to determine where new ML capabilities can meaningfully...  ...high quality agentic search systems, addressing tail latency,...  ...serving, query processing, and ML inference. Turn advances in information...  ...systems, mentor senior engineers, and align technical and product... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  • $117k - $161k

     ...Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the...  ...in on this mission. If you are too, let's talk. The  Senior Systems Engineer Opportunity As a Senior Systems Engineer at Okta, you will... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    10 days ago
  • $147.93k - $291.61k

     ...Description Job Description Waabi, founded by AI visionary Raquel Urtasun, is the leader...  ...Execution: Lead the end-to-end systems engineering lifecycle for the Sensing, Perception, Maps...  ...for leadership. Bonus: - AI/ML Integration: Experience working alongside... 
    Full time
    Contract work
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    7 days ago
  • $70 - $110 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Summers , and Jack Dorsey . Position: Computer and Information Systems Managers Type: Contract Compensation: $70–$110/hour... 
    Hourly pay
    Contract work
    Summer work
    Immediate start
    Remote work

    Mercor

    San Francisco, CA
    3 days ago
  • $172.5k - $313.7k

     ...are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer...  ...startup and the reach of Salesforce. We’re building Systems Integration Agent, announced at Dreamforce, which... 

    100 Salesforce, Inc.

    San Francisco, CA
    23 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI/ML Infra / Systems Engineer (Inference). Be the first to apply!