Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

$300k - $400k

Thinking Machines Lab

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up - powering Tinker's live, multi-tenant serving and the products built on top of our models.

This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.

What You'll Do
  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform
  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production
  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly
  • Partner with inference and research teams to productionize new serving techniques without compromising reliability
  • Lead incident response for production inference issues, driving root cause analysis and durable fixes
  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows
  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale
Skills & Qualifications
Minimum Qualifications
  • Experience operating large-scale, latency-sensitive production systems
  • Proficiency in Python and Go or another systems language
  • Experience with observability, monitoring, and incident response for production services
  • Strong understanding of distributed systems and how they fail at scale
Preferred Qualifications
  • Experience running production inference for large language models or other large-scale ML systems
  • Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags
  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure
  • Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications
  • Comfortable being on-call and leading incident response for critical production systems
  • Comfortable working with high autonomy in a fast-changing, early-stage environment
Logistics
  • Location: This role is based in San Francisco, CA.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in New York, NY vacancy
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference... 
    Suggested
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    3 days ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE...  ...Proficiency in enhancing the performance of software systems, particularly in the context of... 
    Suggested
    Flexible hours

    Jobleads-US

    New York, NY
    15 hours ago
  •  ...Collective Intuition, Inc. seeks a Founding Infrastructure Engineer to define and own the production inference platform behind a new layer of AI intelligence. You will build core systems, set foundational architecture decisions, and influence engineering culture from day... 
    Suggested

    Jobleads-US

    New York, NY
    1 day ago
  •  ...Observable Intuition, Inc. seeks a founding Infrastructure Engineer to define and own the production inference platform behind our data-driven AI layer. You will collaborate with the founding team to build core systems from the ground up, make foundational architectural... 
    Suggested

    Jobleads-US

    New York, NY
    3 days ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction... 
    Suggested
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    2 days ago
  • $155k - $180k

     ...& Growth department is seeking a highly skilled full stack software engineer to support, manage, and improve the AI-driven software the Integrity...  ...outputs into production applications, including deployment, inference, versioning, and monitoring.Integrate third-party and... 
    Odd job
    Full time
    Temporary work
    Local area
    Remote work
    1 day per week

    National Basketball Association

    New York, NY
    2 days ago
  • $190k - $260k

     ...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft...  ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $229.9k - $262.4k

     ...to build world-class applied science and engineering teams to deliver our industry leading capabilities...  ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    5 days ago
  •  ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing...  ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large... 
    Local area

    Capital One

    New York, NY
    1 day ago
  • $225k - $325k

     ...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft...  ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $184.7k - $324.8k

     ...Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference New York City, New York, United States Machine Learning and AI...  ...will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help... 
    Worldwide
    Relocation

    Jobleads-US

    New York, NY
    3 days ago
  •  ...Location: New York, NY We are seeking a Senior Software Engineer to help build the foundational platforms that power enterprise AI products...  ...AI-powered applications, agent-based systems, model-inference services, or AI-serving platforms. Working knowledge of machine... 
    Full time

    PRI Technology

    New York, NY
    3 days ago
  • $130.6k - $192k

     ...platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-...  ...technologies and advanced causal inference and data mining techniques.What We're Looking...  ...Claude Code, Codex, Cursor) across the full software development lifecycle, including design,... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    New York, NY
    1 day ago
  • $193.3k - $261.5k

     ...sales, and more.We are looking for a Senior Software Developer with a passion for dealing...  ...and motivated software and data science engineers who build systems and models used to analyze...  ..., including architecture, training/inference lifecycles, and optimization of model execution... 
    Internship
    Local area
    Flexible hours

    Amazon

    New York, NY
    5 days ago
  • $165k - $250k

     ...About the team Mbodi builds the software layer that lets industrial robots learn new skills from instruction in minutes...  ...distributed agent orchestration, and compiled neural inference. As one of our early engineers, you'll work on the core systems that translate camera... 
    Work at office
    3 days per week

    Mbodi AI

    New York, NY
    3 days ago
  •  ...pace. What you'll do You’ll join a small, talent-dense engineering team with significant ownership across product, infrastructure...  ...js, React, tRPC, Postgres, ClickHouse, Temporal, and AWS. Our inference stack is Python, serving LLMs, TTS, and diffusion image and... 
    Full time
    Work at office

    coreflow

    New York, NY
    1 day ago
  • $158.1k - $213.8k

     ...for building innovation in silicon and software for our AWS customers. We are at the forefront...  ...scale with the world’s most talented engineers. Our team covers multiple disciplines...  ...chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the... 
    Internship
    Flexible hours

    Amazon

    New York, NY
    3 days ago
  • $182k - $242k

     ...March 2025. Learn more at .About the RoleWe are seeking Senior Software Engineers who specialize across the pillars of Observability to play a...  ...AI platforms and workloads (e.g., large-scale training and inference, GPU-based infrastructure, MLOps tooling) is a plus.The base... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    17 hours ago
  • Intone Inc. is seeking a DevOps/Inference Engineer for an 18+ month remote contract to own the infrastructure, self-hosted model serving stack, and platform reliability for a healthcare AI benchmark and evaluation suite targeting clinical prediction tasks such as sepsis... 
    Remote job
    Contract work

    Intone Inc

    New York, NY
    5 days ago
  •  ...Capital One is seeking an AI Engineer 5 in New York to help advance foundation models and LLM inference within the Intelligent Foundations and Experiences team. You will build, optimize, and deploy AI software components across a broad stack, collaborating with engineers... 

    Jobleads-US

    New York, NY
    3 days ago
  • $38.45 - $62.5 per hour

     ...Software Development Engineer in Test (SDET) Location: Remote or Hybrid in TX (Dallas metro-area) Long-Term Contract (3-4 years) Pay Rate: $38....  ...interviewing at ConsultNet Technology Services and Solutions by 2x Inferred from the description for this job Medical insurance... 
    Long term contract
    Contract work
    Work experience placement
    Remote work

    ConsultNet Technology Services and Solutions

    New York, NY
    3 days ago
  • $180k - $220k

     ...how value compounds across the platform. Engineers here treat AI as a force multiplier in...  ...Join an AI-native engineering team as a Software Engineer. The engineering org is currently...  ...that supports production-grade inference, evaluation, and monitoring. • Work across... 
    Full time
    For contractors
    Work at office
    Relocation
    Visa sponsorship

    Syndesus, Inc.

    New York, NY
    5 days ago
  •  ...isn't an AI wrapper slapped onto legacy software - we built a proprietary general ledger...  ...About the role Hanover Park is an engineering-first company on a mission to build the...  ...Trigger.dev for background jobs and AI inference. What we're looking for Need:... 
    Local area

    Hanover Technologies, Inc

    New York, NY
    3 days ago
  •  ...Software Engineer As a Software Engineer, you'll work directly with our Head of Engineering and product team to build the agentic platform...  ...agentic infrastructure that supports production-level inference, evaluation, and monitoring Work across the stack to integrate... 
    Temporary work
    Flexible hours

    Baton, Inc.

    New York, NY
    2 days ago
  • $137.21k - $185.19k

     ...products that bring generative AI into the physical world. As a Software Engineer, you will play a central role in developing the platforms,...  ...span the software-hardware boundary, combining low-latency inference pipelines, robust cloud infrastructure, and tightly... 
    Local area
    Immediate start

    Siemens

    New York, NY
    1 day ago
  •  ...Personalization and Discovery (PVPD) is seeking a Senior Software Development Engineer to join a small, high-caliber team building the next generation...  ...dialogue, and contextual recommendations Optimize LLM inference for latency, cost, and quality at the scale of one of the... 

    Amazon

    New York, NY
    1 day ago
  • Jobgether SRL is seeking an AI Research Engineer (Kernel & Inference Optimization) to advance model-serving architectures for diverse hardware and edge devices. You will blend hands-on research with low-level engineering to push latency, throughput, and memory efficiency... 
    Remote work

    Jobgether SRL

    New York, NY
    4 days ago
  • $102.5k - $210.6k

    Position Summary As an Applied AI Engineer III, you will actively engage in your...  ...engineering craftsmanship across full-stack software engineering and modern frameworks—...  ...designs and implementations, and owning the inference, token, and cloud cost of what you build... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Deloitte

    New York, NY
    1 day ago
  • $225k - $300k

     ...Datalab Fullstack Engineer Salary range: $225k - $300k | Equity: 0.15% - 0.35% | In-Person: NYC About Datalab Datalab trains...  ...and document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks,... 
    Local area

    Data Lab

    New York, NY
    2 days ago
  • $180k - $220k

     ...yr Direct message the job poster from FutureX. Senior Software Engineer (Backend) - Series A AI HealthTech Startup - Hybrid in NYC...  ...increase your chances of interviewing at FutureX. by 2x Inferred from the description for this job Medical insurance Vision... 
    Full time

    Futurex

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!