Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

$300k - $350k
Full-time

Thinking Machines Lab

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up — powering Tinker's live, multi-tenant serving and the products built on top of our models.

This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.

What You'll Do

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform

  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production

  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly

  • Partner with inference and research teams to productionize new serving techniques without compromising reliability

  • Lead incident response for production inference issues, driving root cause analysis and durable fixes

  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows

  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale

Skills & Qualifications

Minimum Qualifications

  • Experience operating large-scale, latency-sensitive production systems

  • Proficiency in Python and Go or another systems language

  • Experience with observability, monitoring, and incident response for production services

  • Strong understanding of distributed systems and how they fail at scale

Preferred Qualifications

  • Experience running production inference for large language models or other large-scale ML systems

  • Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags

  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure

  • Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications

  • Comfortable being on-call and leading incident response for critical production systems

  • Comfortable working with high autonomy in a fast-changing, early-stage environment

Logistics

  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in New York, NY vacancy
  • $155k - $180k

     ...& Growth department is seeking a highly skilled full stack software engineer to support, manage, and improve the AI-driven software the Integrity...  ...outputs into production applications, including deployment, inference, versioning, and monitoring.Integrate third-party and... 
    Suggested
    Odd job
    Full time
    Temporary work
    Local area
    Remote work
    1 day per week

    National Basketball Association

    New York, NY
    1 day ago
  • $229.9k - $262.4k

     ...to build world-class applied science and engineering teams to deliver our industry leading capabilities...  ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  •  ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing...  ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large... 
    Suggested
    Local area

    Capital One

    New York, NY
    5 days ago
  • $190k - $225k

     ...next generation of voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio daily, and power 2...  ...encourage you to apply! About the role: We're hiring a Software Engineer to help turn cutting-edge AI research into products our customers... 
    Suggested

    AssemblyAI

    New York, NY
    8 days ago
  • $193.3k - $261.5k

     ...sales, and more.We are looking for a Senior Software Developer with a passion for dealing...  ...and motivated software and data science engineers who build systems and models used to analyze...  ..., including architecture, training/inference lifecycles, and optimization of model execution... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    New York, NY
    4 days ago
  • $130.6k - $192k

     ...platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-...  ...technologies and advanced causal inference and data mining techniques.What We're Looking...  ...Claude Code, Codex, Cursor) across the full software development lifecycle, including design,... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    New York, NY
    5 hours ago
  • $193.3k - $261.5k

     ...Personalization and Discovery (PVPD) is seeking a Senior Software Development Engineer to join a small, high-caliber team building the next generation...  ...dialogue, and contextual recommendations - Optimize LLM inference for latency, cost, and quality at the scale of one of the... 
    Internship
    Local area
    Flexible hours

    Amazon

    New York, NY
    1 day ago
  •  ...Capital One is seeking an AI Engineer 5 in New York to help advance foundation models and LLM inference within the Intelligent Foundations and Experiences team. You will build, optimize, and deploy AI software components across a broad stack, collaborating with engineers... 

    Jobleads-US

    New York, NY
    1 day ago
  •  ...Software Engineer As a Software Engineer, you'll work directly with our Head of Engineering and product team to build the agentic platform...  ...agentic infrastructure that supports production-level inference, evaluation, and monitoring Work across the stack to integrate... 
    Temporary work
    Flexible hours

    Baton, Inc.

    New York, NY
    1 day ago
  •  ...Nscale seeks a Principal AI Engineer to lead inference and post-training for our GPU cloud, defining multi-year roadmaps and standards. You will guide 20–50+ engineers, shaping architecture from kernel to serving, across disaggregated and BYOM deployments, with emphasis... 

    Jobleads-US

    New York, NY
    1 day ago
  • $185.5k - $232k

     ...Senior Software Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma...  .... ~ Familiarity with ML concepts such as training versus inference, embeddings, retrieval, evaluation, data quality, and model... 
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    New York, NY
    1 day ago
  • $165k - $225k

    Senior Software Engineer, Network Platform Moonlite delivers high-performance AI infrastructure for organizations running intensive computational...  ...networking for distributed computing, model training, inference, and data-intensive workloads. Working closely with our... 
    Immediate start
    Flexible hours

    Moonlite AI

    New York, NY
    3 days ago
  • $160k - $230k

    Senior Software Engineer - Together Cloud Platform San Francisco About the Role Together AI is building the AI Acceleration Cloud, an end-...  ...the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior... 
    Full time
    Remote work

    Together AI

    New York, NY
    3 days ago
  • We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Software Engineer III at JPMorganChase within the Business Banking, Business Access and Tools division, you serve as a seasoned member of an agile team... 

    JP Morgan Chase

    New York, NY
    3 days ago
  • A Software Engineer is needed to design, develop, and maintain modern software applications and services. The engineer will work within a...  ...Docker and Kubernetes. Integrate AI models, data pipelines, and inference services into production systems. Collaborate with cross-... 
    Remote job
    Monday to Friday
    Shift work

    Gnostech

    New York, NY
    3 days ago
  • $130.6k - $192k

    About the TeamCome help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the RoleOur mission is to improve the developer experience of engineers at DoorDash by... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    New York, NY
    3 days ago
  • $38.45 - $62.5 per hour

    Software Development Engineer in Test (SDET) Location: Remote or Hybrid in TX (Dallas metro-area) Long-Term Contract (3-4 years) Pay Rate: $3...  ...interviewing at ConsultNet Technology Services and Solutions by 2x Inferred from the description for this job Medical insurance... 
    Long term contract
    Contract work
    Work experience placement
    Remote work

    ConsultNet Technology Services and Solutions

    New York, NY
    2 days ago
  • Software Engineer Chalk is building the data platform that powers the future of machine learning applications. We tear down complexity, latency...  ...programs in order to optimize arbitrary user Python code, infers and orchestrates infrastructure implied by the structure of that... 
    Work at office
    Flexible hours

    CHALK INC

    New York, NY
    3 days ago
  • $80k - $107k

    GPU Software Engineer (CUDA) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    New York, NY
    5 days ago
  • GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable... 
    Flexible hours

    Baseten

    New York, NY
    3 days ago
  •  ...GitHub Actions), K8s, managed Postgres, and AI inference. Today, we serve 500+ customers on our managed cloud. We're now hiring engineers to help us with our Postgres, GitHub...  ...tools / techniques more and more during our software development processes. We'd like to share... 
    Immediate start

    Ubicloud

    New York, NY
    3 days ago
  • Software Engineer II Technology is at the heart of Disney's past, present, and future. Disney Entertainment and ESPN Product & Technology...  ...time-series data, model training on GPU clusters, real-time inference pipelines, and model improvement. You will partner with engineering... 

    Disney

    New York, NY
    3 days ago
  •  ...how value compounds across the platform. Engineers here use AI as a force multiplier in...  ...moments of their lives. About the Role As a Software Engineer, you'll work directly with our...  ...that supports production-level inference, evaluation, and monitoring Work across... 
    Temporary work
    Work at office
    Flexible hours

    Baton Market Inc

    New York, NY
    3 days ago
  • $75.5 - $102 per hour

     ...make a profound impact, empowering every engineering team with safe, governed, and cost-...  ...-7 years of professional experience in software development. Strong proficiency in Python...  ...model routing, semantic caching, and batch inference. Experience building developer... 
    Hourly pay
    Temporary work
    Flexible hours

    Aquent

    New York, NY
    2 days ago
  •  ...an exciting and rewarding opportunity for you to take your software engineering career to the next level. The Chief Data & Analytics Office...  ...secure and high-quality production code for machine learning inference and training systemsProduces architecture and design... 
    Work at office

    JP Morgan Chase

    New York, NY
    1 day ago
  • $160k - $240k

    Senior Software Engineer - RDF Infrastructure Location New York Business Area Engineering and CTO Ref # 10054258 Description...  ...isolation semantics, statistics and cardinality estimation, inference, replication, scalability, and operational reliability.The... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    3 days ago
  • $132.6k - $192.3k

     ...for this position. Skills and Competencies ~5+ years of software engineering experience designing, coding, testing, and operating...  ...services, application programming interfaces, data pipelines, and inference pipelines that support real-time and batch AI workloads... 
    Full time

    Moody's

    New York, NY
    a month ago
  •  ...Nscale is seeking a Principal AI Engineer to lead the inference and post-training pillar of our AI systems engineering organization in the United States. You will shape the multi-year technical roadmap for serving models and post-training workloads, guiding 20–50+ engineers... 

    Jobleads-US

    New York, NY
    1 day ago
  • $225k - $300k

     ...training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction...  ...document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks, and... 
    Local area

    datalab.mn d.o.o.

    New York, NY
    1 day ago
  • $91.7k - $163.7k

    Sr Software Engineer Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives...  ..., metadata enrichment, and performance tuning for scalable inference Partner with broader analytics and AI teams to make recommendations... 
    Minimum wage
    Full time
    Work experience placement
    Local area
    Remote work

    UnitedHealthcare At Home

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!