Get new jobs by email
  • $8k

     ...TS/SCI) clearance with polygraph is required.  Visionist has an exciting new, fully FUNDED opportunity for a Software Engineer - Inference on our largest PRIME contract. Our team of Analysts and Engineers is motivated by the direct impact on the mission, crafting... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Temporary work
    Immediate start
    Flexible hours

    Visionist, Inc.

    Remote
    1 day ago
  • $2,000 per month

     ...the Etched serving front-end. Representative projects Build scheduling logic for handling continuous batching and real time inference Implement inference-time acceleration techniques such as speculative decoding, tree search, KV cache sharing, etc. Implement... 
    Suggested
    Full time
    Work at office
    Relocation package

    Etched

    Remote
    1 day ago
  • $160k - $230k

     ...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and... 
    Suggested
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  •  ...the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is... 
    Suggested
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $188k - $275k

     ...became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  . What You'll Do Description of the team: The Inference team is responsible for delivering high-performance model serving capabilities that meet the needs of real production workloads.... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Core Weave

    Remote
    1 day ago
  •  ...P-1284 About This Role As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language... 
    Suggested
    Full time

    Databricks

    Remote
    1 day ago
  •  ...the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is... 
    Suggested
    Full time

    Cerebras Systems

    United States
    1 day ago
  • $188k - $275k

     ...Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  . What You’ll Do: Inference Platform Team The Inference team builds and operates CoreWeave’s Kubernetes-native inference platform, powering low-latency, high... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    1 day ago
  • $320k

     ...policy experts, and business leaders working together to build beneficial AI systems. About the Role Our mandate is to make inference deployment boring and unattended. Anthropic serves Claude to millions of users across GPUs, TPUs, and Trainium — and every... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model... 
    Suggested
    Full time
    Work at office
    3 days per week

    Pika

    Remote
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $96.4k - $241k

    Global BiostatisticsJob Level: Statistical ScientistLocation: Home-based in the U.S / CanadaJoin us on our exciting journey!The Global...  ...therapeutic areas. IQVIA Biostatistics helps interpret and draw inferences from data collected on patients as they progress through a... 
    Suggested
    Full time
    Part time
    Immediate start
    Work from home
    Worldwide

    IQVIA

    Durham, NC
    4 hours ago
  • $135k - $160k

     ...the technologies to make this possible, with the ultimate goal of enabling human life on Mars. APPLICATION SOFTWARE ENGINEER, INFERENCE The application software team is the central nervous system of SpaceX – we create mission critical applications that are used throughout... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Remote work
    Worldwide
    Weekend work

    Spacex

    Remote
    1 day ago
  •  ...characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference. ~ Strong understanding or working experience with distributed systems. ~ Experience in Golang, C++ or other languages designed... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $110k - $150k

     ...production and helping with GPU/CPU optimization. You’ll also work with various hardware partners (NVIDIA, AMD, Intel, Apple) to optimize inference on their hardware. About you Hands-on experience with performance optimization, e.g. concurrency, multithreading, memory,... 
    Suggested
    Full time
    Work experience placement
    Relocation

    Topaz Labs

    Dallas, TX
    1 day ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is... 
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  •  ...their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment... 
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    4 hours ago
  • $300k

     ...engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    1 day ago
  • $320k

     ...policy experts, and business leaders working together to build beneficial AI systems. About the Role Our mandate is to make inference deployment boring and unattended. Anthropic serves Claude to millions of users across GPUs, TPUs, and Trainium — and every... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    Remote
    1 day ago
  •  ...Title:  Applied AI Engineer — Inference & Agent Systems Location: United States What We're Building Arcana is building AI agents that synthesize information across heterogeneous sources and deliver structured, reasoned answers in real time. The product only works... 
    Full time

    Arcana Analytics

    United States
    1 day ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $320k

     ...engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Remote
    1 day ago
  • $200k - $250k

     ...to-end lifecycle of our ML systems — from packaging and deployment to monitoring, performance, and scaling — across a custom-built inference platform powering a live conversational product. This isn’t a typical “pipeline” role. Our platform runs multiple specialized... 
    Full time
    Remote work
    Flexible hours

    Wizard

    United States
    1 day ago
  • $170k - $216k

     ...hear from you! In this hybrid role you will report to the Software Engineering Manager.   You will: Build and evolve ML inference infrastructure for simulations. Be responsible for the reliability, latency, and user experience of ML model deployment and... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprises and developers alike to use and access our state-of-the-art AI models, allowing them to do things that they’ve... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • Job-ID27995022Reference24-01027Responsibilities and Requirements:Experience working in various disease areas.They also must have specific expertise in SDTM programming, have submissions experience, and MUST possess strong skill in graphing.They need to be comfortable working...

    Katalyst Healthcares & Life Sciences

    Trumbull, CT
    4 hours ago