Software Engineer, Inference
$300k - $400kThinking Machines Lab
About Thinking Machines The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. About the Role We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up - powering Tinker's live, multi-tenant serving and the products built on top of our models. This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online. What You'll Do
Minimum Qualifications
- Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform
- Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production
- Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly
- Partner with inference and research teams to productionize new serving techniques without compromising reliability
- Lead incident response for production inference issues, driving root cause analysis and durable fixes
- Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows
- Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale
Minimum Qualifications
- Experience operating large-scale, latency-sensitive production systems
- Proficiency in Python and Go or another systems language
- Experience with observability, monitoring, and incident response for production services
- Strong understanding of distributed systems and how they fail at scale
- Experience running production inference for large language models or other large-scale ML systems
- Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags
- Experience with capacity planning and cost optimization for GPU or TPU infrastructure
- Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications
- Comfortable being on-call and leading incident response for critical production systems
- Comfortable working with high autonomy in a fast-changing, early-stage environment
- Location: This role is based in San Francisco, CA.
- Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.
- Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in New York, NY vacancy
$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference...SuggestedTemporary workFor contractorsWork experience placement- ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI... ...Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE... ...Proficiency in enhancing the performance of software systems, particularly in the context of...SuggestedFlexible hours
- ...Collective Intuition, Inc. seeks a Founding Infrastructure Engineer to define and own the production inference platform behind a new layer of AI intelligence. You will build core systems, set foundational architecture decisions, and influence engineering culture from day...Suggested
- ...Observable Intuition, Inc. seeks a founding Infrastructure Engineer to define and own the production inference platform behind our data-driven AI layer. You will collaborate with the founding team to build core systems from the ground up, make foundational architectural...Suggested
$250k - $300k
Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction...SuggestedWork experience placementWork at officeLocal areaImmediate start$155k - $180k
...& Growth department is seeking a highly skilled full stack software engineer to support, manage, and improve the AI-driven software the Integrity... ...outputs into production applications, including deployment, inference, versioning, and monitoring.Integrate third-party and...Odd jobFull timeTemporary workLocal areaRemote work1 day per week$190k - $260k
...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft... ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems...Full timeWork experience placementWork at officeLocal areaRemote workHome office$229.9k - $262.4k
...to build world-class applied science and engineering teams to deliver our industry leading capabilities... ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model...Full timePart timeLocal area- ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing... ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large...Local area
$225k - $325k
...they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft... ...), especially how they influence latency and throughput of inference.Strong understanding or working experience with distributed systems...Full timeWork experience placementWork at officeLocal areaRemote workHome office$184.7k - $324.8k
...Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference New York City, New York, United States Machine Learning and AI... ...will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help...WorldwideRelocation- ...Location: New York, NY We are seeking a Senior Software Engineer to help build the foundational platforms that power enterprise AI products... ...AI-powered applications, agent-based systems, model-inference services, or AI-serving platforms. Working knowledge of machine...Full time
$130.6k - $192k
...platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-... ...technologies and advanced causal inference and data mining techniques.What We're Looking... ...Claude Code, Codex, Cursor) across the full software development lifecycle, including design,...Hourly payWork at officeLocal areaRemote workFlexible hours$193.3k - $261.5k
...sales, and more.We are looking for a Senior Software Developer with a passion for dealing... ...and motivated software and data science engineers who build systems and models used to analyze... ..., including architecture, training/inference lifecycles, and optimization of model execution...InternshipLocal areaFlexible hours$165k - $250k
...About the team Mbodi builds the software layer that lets industrial robots learn new skills from instruction in minutes... ...distributed agent orchestration, and compiled neural inference. As one of our early engineers, you'll work on the core systems that translate camera...Work at office3 days per week- ...pace. What you'll do You’ll join a small, talent-dense engineering team with significant ownership across product, infrastructure... ...js, React, tRPC, Postgres, ClickHouse, Temporal, and AWS. Our inference stack is Python, serving LLMs, TTS, and diffusion image and...Full timeWork at office
$158.1k - $213.8k
...for building innovation in silicon and software for our AWS customers. We are at the forefront... ...scale with the world’s most talented engineers. Our team covers multiple disciplines... ...chips. Inferentia delivers best-in-class ML inference performance at the lowest cost in the...InternshipFlexible hours$182k - $242k
...March 2025. Learn more at .About the RoleWe are seeking Senior Software Engineers who specialize across the pillars of Observability to play a... ...AI platforms and workloads (e.g., large-scale training and inference, GPU-based infrastructure, MLOps tooling) is a plus.The base...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- Intone Inc. is seeking a DevOps/Inference Engineer for an 18+ month remote contract to own the infrastructure, self-hosted model serving stack, and platform reliability for a healthcare AI benchmark and evaluation suite targeting clinical prediction tasks such as sepsis...Remote jobContract work
- ...Capital One is seeking an AI Engineer 5 in New York to help advance foundation models and LLM inference within the Intelligent Foundations and Experiences team. You will build, optimize, and deploy AI software components across a broad stack, collaborating with engineers...
$38.45 - $62.5 per hour
...Software Development Engineer in Test (SDET) Location: Remote or Hybrid in TX (Dallas metro-area) Long-Term Contract (3-4 years) Pay Rate: $38.... ...interviewing at ConsultNet Technology Services and Solutions by 2x Inferred from the description for this job Medical insurance...Long term contractContract workWork experience placementRemote work$180k - $220k
...how value compounds across the platform. Engineers here treat AI as a force multiplier in... ...Join an AI-native engineering team as a Software Engineer. The engineering org is currently... ...that supports production-grade inference, evaluation, and monitoring. • Work across...Full timeFor contractorsWork at officeRelocationVisa sponsorship- ...isn't an AI wrapper slapped onto legacy software - we built a proprietary general ledger... ...About the role Hanover Park is an engineering-first company on a mission to build the... ...Trigger.dev for background jobs and AI inference. What we're looking for Need:...Local area
- ...Software Engineer As a Software Engineer, you'll work directly with our Head of Engineering and product team to build the agentic platform... ...agentic infrastructure that supports production-level inference, evaluation, and monitoring Work across the stack to integrate...Temporary workFlexible hours
$137.21k - $185.19k
...products that bring generative AI into the physical world. As a Software Engineer, you will play a central role in developing the platforms,... ...span the software-hardware boundary, combining low-latency inference pipelines, robust cloud infrastructure, and tightly...Local areaImmediate start- ...Personalization and Discovery (PVPD) is seeking a Senior Software Development Engineer to join a small, high-caliber team building the next generation... ...dialogue, and contextual recommendations Optimize LLM inference for latency, cost, and quality at the scale of one of the...
- Jobgether SRL is seeking an AI Research Engineer (Kernel & Inference Optimization) to advance model-serving architectures for diverse hardware and edge devices. You will blend hands-on research with low-level engineering to push latency, throughput, and memory efficiency...Remote work
$102.5k - $210.6k
Position Summary As an Applied AI Engineer III, you will actively engage in your... ...engineering craftsmanship across full-stack software engineering and modern frameworks—... ...designs and implementations, and owning the inference, token, and cloud cost of what you build...Work at officeLocal areaVisa sponsorshipFlexible hours$225k - $300k
...Datalab Fullstack Engineer Salary range: $225k - $300k | Equity: 0.15% - 0.35% | In-Person: NYC About Datalab Datalab trains... ...and document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks,...Local area$180k - $220k
...yr Direct message the job poster from FutureX. Senior Software Engineer (Backend) - Series A AI HealthTech Startup - Hybrid in NYC... ...increase your chances of interviewing at FutureX. by 2x Inferred from the description for this job Medical insurance Vision...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!
Related searches
- software engineer full time New York, NY
- graduate software developer no experience New York, NY
- software support engineer New York, NY
- software developer apprenticeship New York, NY
- software engineer healthcare New York, NY
- network software engineer New York, NY
- software engineer co-op New York, NY
- software engineer internship New York, NY
- senior software engineer New York, NY
- software system engineer New York, NY



