Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI/ML Infra / Systems Engineer (Inference)

$204k - $216k

Sapience AI Corporation

Job Description

Job Description

About Sapience AI

Sapience AI is the collective intelligence platform for professional communities. We sit above the CRMs, AMS platforms, and knowledge bases that organizations already run, and we turn the expertise scattered across them into something every member can search, act on, and share.

The intelligence a community needs is already inside it. Most organizations just cannot reach it. Knowledge lives in silos, in legacy systems, in the heads of a few experts, and in fragmented records no one can connect. We change that.

Our work is grounded in four commitments: technology elevates people and never replaces them, the best expertise is already inside the community, everything is built on trust, and every deployment is purpose-driven for the organization it serves.

Let's achieve more, together.

Where this role sits

This role owns how Sapience AI serves intelligence at scale. You build and run the inference infrastructure that turns the models and reasoning behind MINERVA and COGENT into fast, reliable, affordable answers for members.

You work where models meet production: serving, scaling, latency, and cost. Every time a member asks the platform something, your systems are what make the answer arrive quickly and hold up under load.

You sit close to applied AI, research, and platform engineering, and you make the difference between a model that works in a notebook and one that serves a community reliably.

Why this role exists

Collective intelligence is only useful if it is fast and dependable. Members will not wait, and communities cannot rely on a platform that buckles under load or costs too much to run.

Serving modern AI at scale is a hard systems problem: large models, tight latency budgets, expensive hardware, and demand that spikes. It takes real infrastructure engineering to get right.

The AI/ML Infrastructure and Systems Engineer owns that problem. You make inference fast, reliable, and affordable, so the intelligence the platform promises actually reaches members.

What you will own (Areas of Responsibility)

You hold seven areas of responsibility across inference infrastructure. Each one is yours to set direction on, build, and measure.

1. Inference serving and systems
  • Build and operate the inference systems that serve models and reasoning behind MINERVA and COGENT.
  • Own the serving path end to end, from request to response, under real load.
  • Make serving robust to failure, traffic spikes, and change.
2. Latency and performance
  • Drive down latency so members get fast answers, and keep it low as the platform grows.
  • Optimize the full path, including model execution, batching, caching, and retrieval.
  • Profile relentlessly and remove the bottlenecks that matter.
3. Scale and reliability
  • Scale inference to more members, more communities, and heavier reasoning without losing reliability.
  • Build autoscaling, load management, and graceful degradation.
  • Own the reliability of the serving layer as a first-order responsibility.
4. Cost and efficiency
  • Own the economics of inference, including GPU and hardware efficiency at scale.
  • Improve utilization and cut waste without hurting quality or speed.
  • Make cost a designed property, not a surprise.
5. Model deployment and lifecycle
  • Build the paths that take models and reasoning components from research to production safely.
  • Support rollout, versioning, and rollback with confidence.
  • Give applied AI and research a fast, safe route to ship.
6. Observability and operations
  • Instrument the serving layer so it can be measured, debugged, and trusted.
  • Build the alerts, dashboards, and tooling that keep inference healthy.
  • Reduce the operational burden of running AI at scale.
7. Hardware, accelerators, and platform choices
  • Make sound choices about accelerators, runtimes, and serving frameworks.
  • Balance performance, cost, and maintainability in platform decisions.
  • Keep the stack current as inference technology moves.
AI-augmented ways of working

You build the infrastructure that serves AI, and you use AI in building it, to generate tooling, reason about performance, and move faster, while you own correctness, reliability, and cost.

The standard is human in partnership: AI accelerates the work, you own the judgment, the interpretation, and the call. The people who create the most value here are not the ones producing the most output. They are the ones turning evidence into clear, durable decisions.

What this role is not

To keep the boundary clear:

  • This is not a model research role. You serve and scale models; you do not develop new model architectures.
  • This is not a general backend role. Your center of gravity is inference systems, performance, and hardware efficiency.
  • This is not a data engineering role. You partner with data teams, but you own serving, not the data platform.
  • This is not a best-effort prototype role. You are accountable for production inference members depend on.
What success looks like

We measure this role on outcomes the team can see:

  • Fast answers. Members get low-latency responses, and latency stays low as the platform grows.
  • Reliable at scale. Inference holds up under load, spikes, and growth.
  • Affordable intelligence. Inference cost per answer improves as usage rises.
  • Safe rollouts. Models and reasoning components ship, version, and roll back with confidence.
  • Operable serving. The serving layer is instrumented, debuggable, and healthy.
  • Sound platform choices. Accelerator and framework decisions age well.
Who you are Required qualifications
  • Five or more years in infrastructure, systems, or ML infrastructure engineering.
  • Hands-on experience serving ML or LLM models in production at scale.
  • Deep understanding of latency, throughput, and performance optimization.
  • Experience with GPUs or accelerators and their efficient use.
  • Strong systems programming and distributed systems fundamentals.
  • A track record of reliable, cost-aware production systems.
  • Fluency with observability and operational excellence.
Preferred qualifications
  • Experience with inference-serving frameworks and model runtimes.
  • Experience optimizing LLM inference, including batching, quantization, and caching.
  • Familiarity with retrieval systems and their performance characteristics.
  • Experience owning cost and capacity for AI workloads.
  • Exposure to neuro-symbolic or agentic systems in production.
How you work
  • You name the real bottleneck before reaching for a fix.
  • You measure before and after, and you trust evidence over intuition.
  • You treat reliability and cost as first-order, alongside speed.
  • You build systems others can operate.
  • You share tooling and knowledge across the team.
Skills & Competencies
  • Inference serving architecture and optimization.
  • Latency, throughput, and performance engineering.
  • GPU and accelerator efficiency.
  • Distributed systems, scaling, and reliability engineering.
  • Model deployment, versioning, and rollout.
  • Observability and production operations for AI.
  • Cost and capacity management for inference.
Services & Tools Experience
  • Inference-serving frameworks and model runtimes (for example vLLM, TensorRT, Triton-class systems).
  • GPU tooling, CUDA-class ecosystems, and accelerator runtimes.
  • Kubernetes, containers, and cloud platforms (AWS, GCP, or Azure).
  • Observability stacks (metrics, tracing, and logging).
  • Python and a systems language such as Go, Rust, or C++.
  • Caching, queuing, and load-management systems.
  • Serving the models and reasoning behind the MINERVA platform and COGENT architecture.
Prior Experience & Background
  • Prior ML infrastructure, platform, or systems engineering at a software or AI company.
  • Experience serving models in production under real load.
  • A track record of performance and reliability improvements at scale.
  • Experience owning inference cost is a plus.
Cross-functional partners

You work most closely with Applied AI, Neuro-Symbolic AI, Platform Engineering, and Research. You serve the models and reasoning that power MINERVA and the COGENT architecture, and you own how they run in production.

How we hire

We review every application, and we encourage you to apply even if you do not match every line above. Research shows that talented people, especially those from underrepresented communities, often hold back when they do not meet every qualification. If that is the only thing holding you back, apply anyway.

Sapience AI is an equal opportunity employer. We are committed to a workplace where everyone, regardless of background, has a voice in building what comes next.

Compensation

Base Salary:  $204,000 - $216,000 + early stage equity

Generous health and wellness benefits

 

Sapience AI is an equal opportunity employer. We do not discriminate on the basis of gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. If you need an accommodation to complete our application process, let your recruiter know.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI/ML Infra / Systems Engineer (Inference) in Los Angeles, CA vacancy
  • $135k - $220k

     ...to reindustrialize America. By combining AI, advanced software, robotics, and full-...  ...aircraft, ships, and other mission-critical systems up to 10x faster and at significantly...  ...the Machine Controls & Data Integration Engineer.AI/ML Model Development: Design, develop, and... 
    Suggested
    Permanent employment
    Full time
    Relocation package
    Flexible hours

    Hadrian

    Los Angeles, CA
    4 days ago
  • $102.5k - $187.9k

     ...organizations. Using our product-driven, AI-centric approach, we empower organizations...  ...data scientists, designers, and software engineers enable our clients to solve their most...  ...agile delivery process Work on distributed systems problems ranging from scheduling,... 
    Suggested
    Summer holiday
    Flexible hours

    EY

    Los Angeles, CA
    1 day ago
  • $124k - $186k

     ...Applied Intelligence Data Engineering team seeks a Senior...  ...real-time data streaming systems. It builds high-...  ...-time analytics, APIs, AI workflows, and mission-...  ...Collaborate with AI/ML teams to enable real-time...  ...feature pipelines and inference services. Technical... 
    Suggested
    Burbank, CA
    more than 2 months ago
  •  ...Senior Systems Engineer VAST Data is looking for a Senior Systems Engineer to join our growing...  ...Data is the data platform company for the AI era. We are building the enterprise...  ...-time data analysis and AI training and inference. Designed from the ground up to make AI... 
    Suggested
    Traineeship

    VAST Data

    Los Angeles, CA
    4 days ago
  •  ...and fabricate seamless, custom-carved wall systems that live at the intersection of...  ..., and bespoke design. We're building an AI-native company and already use AI across...  ...workflows. We're looking for an AI Systems Engineer to help us build, improve, and maintain AI... 
    Suggested
    Full time
    Part time

    MR Walls

    Santa Monica, CA
    1 day ago
  • $77k - $202k

    The Opportunity As a GenAI Python Systems Engineer - Senior Associate, you will play a pivotal role in transforming raw data into actionable...  ..., you will focus on developing and implementing advanced AI and ML solutions to drive innovation and enhance business processes... 
    H1b

    PwC

    Los Angeles, CA
    3 days ago
  • NEOTech is seeking a Director, Business Systems & Manufacturing Applications to lead our internal systems strategy across ERP, MES, BI and AI-enabled solutions. You will own the roadmap, governance, and vendor accountability, partnering with IT Executive Management to standardize... 

    NEOTech

    Los Angeles, CA
    2 days ago
  • $165k - $265k

     ...human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our...  ..., test, and operate all parts of the system - receivers that allow users to connect within...  ..., validate, and productize solutions for AI clusters (100k+ GPU scale)Develop... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $113k - $210k

     ...we hiring? Sphere is seeking a Senior Systems Development Engineer to design, build, automate, and operate...  ...qualifications ~ Experience building or operating AI/ML infrastructure, including GPU compute, model training or inference, data pipelines, and workload scheduling.... 
    Full time
    Local area
    Remote work

    Sphere Entertainment Group, LLC

    Burbank, CA
    4 days ago
  • $250k - $360k

    Join to apply for the Staff Backend Engineer (AI & Knowledge Systems) role at Demand.io Base Pay Range $250,000.00/yr - $360,000.00/yr Additional Compensation Types Annual Bonus and Stock options The Opportunity Demand.io is a profitable, founder‑led consumer AI commerce... 
    Full time

    Demand.io

    Santa Monica, CA
    5 days ago
  • $175k - $308.5k

     ...RF Modeling Systems Engineer At Apple, new ideas have a way of becoming extraordinary products and customer experiences. Bring passion...  ...circuit designers to capture designs and develop models. - Develop AI/ML modem solutions and mitigation techniques.... 
    Relocation

    Apple

    Los Angeles, CA
    5 days ago
  • $141.4k - $234.5k

     ...companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and...  ...space, is seeking a talented and motivated Senior Pre-Sales Systems Engineer within a Pod model for an emerging business positioned for high... 
    Remote work
    Flexible hours

    Pure Storage

    Los Angeles, CA
    1 day ago
  • $99.2k - $138.88k

     ...reusable, safe, and low-cost space vehicles and systems within a culture of safety, collaboration...  ...part of Advanced Concepts and Enterprise Engineering (ACE), supporting Blue Origin's mission...  ..., and uncertainty. Utilize advanced AI-assisted engineering and software-... 
    Permanent employment
    Temporary work
    Local area
    Relocation

    Blue Origin

    Los Angeles, CA
    4 days ago
  • $87.5k - $131.1k

     ...environments. Implement automation for system provisioning, configuration management, and monitoring. Collaborate with engineering, DevOps, and InfoSec teams to ensure secure...  ...such as Git Ops, Ansible, Terraform, Claude AI. ~ Familiarity with container orchestration... 
    Full time
    Work at office
    Remote work
    3 days per week

    Green Dot

    Los Angeles, CA
    4 days ago
  • $100k - $175k

     ...ability to produce nuclear fuel to power AI, advanced manufacturing and critical industries...  .... Our lean, world-class team of engineers and operators is applying a first-...  ...About This Role We are seeking a Fluid Systems Engineer to lead the design, analysis,... 
    Full time
    Weekend work

    General Matter

    Los Angeles, CA
    2 days ago
  • $45 per hour

     ...Senior Systems Engineer We are seeking a skilled and service-oriented Senior Systems Engineer to join our team. This role is primarily focused...  ...as part of infrastructure projects Support and help guide AI enablement initiatives within client environments Configure... 
    Remote work

    DCG Technical Solutions Inc

    Los Angeles, CA
    1 day ago
  • $138.55k - $187.45k

     ...hardware iteration, we deliver high-speed systems at the pace of the modern battlefield. We...  ...About The Role: As an RF systems engineer you will be responsible for identifying RF...  ...3 We may use artificial intelligence (AI) tools to support parts of the hiring process... 
    Weekly pay
    Permanent employment
    Work at office

    Hermeus

    Los Angeles, CA
    3 days ago
  •  ...access control integrity. Focus on anticipating, managing and mitigating outsider and insider threats to company infrastructure, systems. Streamline procedures through automation project management servers. Stand-up, deploy, maintain and monitor servers, metal... 
    Permanent employment
    Local area
    Remote work

    Lucid Circuit, Inc.

    Santa Monica, CA
    3 days ago
  •  ...Systems Engineer Mission LENS’s mission is to end traffic accidents. LENS prevents accidents...  ...: CHP LA Communications Center using AI to prevent accidents before they happen...  ...and physical hardware Experience with ML model optimization and metadata schema... 
    Local area

    RiseMe

    Los Angeles, CA
    3 days ago
  • ZipRecruiter is seeking a Senior BI Engineer to design data systems, models, and AI-powered tooling at the intersection of data, business, and engineering. You will partner with RevOps, Decision Science, Product, and Engineering to scale analysis and decision-making. You... 

    ZipRecruiter, Inc.

    Santa Monica, CA
    5 days ago
  • $52 - $72 per hour

     ...Job Title: Model Based Systems EngineerJob Description The Model-Based Systems Engineer (MBSE) will support the development of advanced hypersonic wind tunnel capabilities...  ...liability. Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to... 
    Permanent employment
    Contract work
    Temporary work
    Work at office
    Work from home
    2 days per week
    3 days per week

    Actalent

    Los Angeles, CA
    4 days ago
  • CommVault Systems Engineer (Data Protection / Backup) Employment Type: Full-Time, Experienced Department: Technology Support CGS is seeking...  ...Email: ****@*****.*** We may use artificial intelligence (AI) tools to support parts of the hiring process, such as... 
    Full time
    Flexible hours

    CGS Federal (Contact Government Services)

    Los Angeles, CA
    3 days ago
  •  ...operates through agile, convergent action teams that integrate engineering, science, and systems analysis. These teams have contributed to award‑winning...  ...spin‑out companies such as CarbonBuilt Inc., Concrete‑AI Inc., and Equatic Inc. Position Summary ICM is seeking a... 

    University of California, Los Angeles

    Los Angeles, CA
    2 days ago
  • $150k - $220k

     ...to reindustrialize America. By combining AI, advanced software, robotics, and full-stack...  ..., ships, and other mission-critical systems up to 10x faster and at significantly lower...  ...We’re looking for a hands-on systems engineer who can take metal additive manufacturing... 
    Permanent employment
    Full time
    Relocation package
    Flexible hours

    Hadrian Automation

    Los Angeles, CA
    9 days ago
  •  ...Vice President of Systems Engineering About the Company Innovative digital banking provider for credit unions & community banks Industry...  ...51-200 Specialties banking api conversational ai online banking online account opening online loan origination... 
    Shift work

    Confidential

    Los Angeles, CA
    3 days ago
  • $139.2k - $208.8k

     ...Overview We are looking for a Senior MLOps Engineer (IC3) to join the Platform Engineering pod. Your mission is to build the Agentic Systems Platform—the runtime environment where...  ...you will bridge the gap between traditional ML serving and autonomous agency. You will... 
    Shift work
    Burbank, CA
    more than 2 months ago
  • $138k - $167k

     ...virtualization technologies and Operating Systems (i.e., Windows, Linux, etc.)...  ...drawing tools Contextual knowledge of AI technologies and it’s practical implementations...  ...real world and how it applies to system engineering and operations Project management: Waterfall... 

    Sony Pictures Entertainment

    Culver City, CA
    1 day ago
  • Systems EngineerPosition OverviewWe are seeking an experienced Systems Engineer to design, deploy, and maintain reliable, secure, and scalable infrastructure across on-premises and cloud environments. The ideal candidate will combine strong systems and networking fundamentals... 

    CyberCoders

    Pasadena, CA
    3 days ago
  •  ..., transmitted and delivered as global energy demands grow. From massive data centers to modernizing transmission systems, our industry-recognized engineers and scientists have been at the forefront of grid transformation for more than a century. You’ll work side-by-side... 
    Work at office
    Local area

    HDR

    Los Angeles, CA
    1 day ago
  •  ...Job Description Job Description Launch Your Engineering Career with Birdi Systems!   Birdi Systems, Inc. (BSI) is seeking a motivated Junior Systems Engineer  with a background in Electrical, Electronic, Mechanical, or Systems Engineering to support, maintain,... 
    Internship
    Work at office

    Birdi Systems, Inc.

    South Pasadena, CA
    9 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI/ML Infra / Systems Engineer (Inference). Be the first to apply!