Software Engineer, Inference
$300k - $350kThinking Machines Lab
About Thinking Machines
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
About the Role
We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up — powering Tinker's live, multi-tenant serving and the products built on top of our models.
This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.
What You'll Do
Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform
Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production
Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly
Partner with inference and research teams to productionize new serving techniques without compromising reliability
Lead incident response for production inference issues, driving root cause analysis and durable fixes
Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows
Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale
Skills & Qualifications
Minimum Qualifications
Experience operating large-scale, latency-sensitive production systems
Proficiency in Python and Go or another systems language
Experience with observability, monitoring, and incident response for production services
Strong understanding of distributed systems and how they fail at scale
Preferred Qualifications
Experience running production inference for large language models or other large-scale ML systems
Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags
Experience with capacity planning and cost optimization for GPU or TPU infrastructure
Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications
Comfortable being on-call and leading incident response for critical production systems
Comfortable working with high autonomy in a fast-changing, early-stage environment
Logistics
Location: This role is based in San Francisco, CA.
Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.
Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
$155k - $180k
...& Growth department is seeking a highly skilled full stack software engineer to support, manage, and improve the AI-driven software the Integrity... ...outputs into production applications, including deployment, inference, versioning, and monitoring.Integrate third-party and...SuggestedOdd jobFull timeTemporary workLocal areaRemote work1 day per week$229.9k - $262.4k
...to build world-class applied science and engineering teams to deliver our industry leading capabilities... ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model...SuggestedFull timePart timeLocal area- ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing... ...Capital One. Design, develop, test, deploy, and support AI software components including foundation model training, large...SuggestedLocal area
$190k - $225k
...next generation of voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio daily, and power 2... ...encourage you to apply! About the role: We're hiring a Software Engineer to help turn cutting-edge AI research into products our customers...Suggested$193.3k - $261.5k
...sales, and more.We are looking for a Senior Software Developer with a passion for dealing... ...and motivated software and data science engineers who build systems and models used to analyze... ..., including architecture, training/inference lifecycles, and optimization of model execution...SuggestedInternshipLocal areaFlexible hours$130.6k - $192k
...platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-... ...technologies and advanced causal inference and data mining techniques.What We're Looking... ...Claude Code, Codex, Cursor) across the full software development lifecycle, including design,...Hourly payWork at officeLocal areaRemote workFlexible hours$193.3k - $261.5k
...Personalization and Discovery (PVPD) is seeking a Senior Software Development Engineer to join a small, high-caliber team building the next generation... ...dialogue, and contextual recommendations - Optimize LLM inference for latency, cost, and quality at the scale of one of the...InternshipLocal areaFlexible hours- ...Capital One is seeking an AI Engineer 5 in New York to help advance foundation models and LLM inference within the Intelligent Foundations and Experiences team. You will build, optimize, and deploy AI software components across a broad stack, collaborating with engineers...
- ...Software Engineer As a Software Engineer, you'll work directly with our Head of Engineering and product team to build the agentic platform... ...agentic infrastructure that supports production-level inference, evaluation, and monitoring Work across the stack to integrate...Temporary workFlexible hours
- ...Nscale seeks a Principal AI Engineer to lead inference and post-training for our GPU cloud, defining multi-year roadmaps and standards. You will guide 20–50+ engineers, shaping architecture from kernel to serving, across disaggregated and BYOM deployments, with emphasis...
$185.5k - $232k
...Senior Software Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma... .... ~ Familiarity with ML concepts such as training versus inference, embeddings, retrieval, evaluation, data quality, and model...Work at officeLocal areaRelocation3 days per week$165k - $225k
Senior Software Engineer, Network Platform Moonlite delivers high-performance AI infrastructure for organizations running intensive computational... ...networking for distributed computing, model training, inference, and data-intensive workloads. Working closely with our...Immediate startFlexible hours$160k - $230k
Senior Software Engineer - Together Cloud Platform San Francisco About the Role Together AI is building the AI Acceleration Cloud, an end-... ...the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior...Full timeRemote work- We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Software Engineer III at JPMorganChase within the Business Banking, Business Access and Tools division, you serve as a seasoned member of an agile team...
- A Software Engineer is needed to design, develop, and maintain modern software applications and services. The engineer will work within a... ...Docker and Kubernetes. Integrate AI models, data pipelines, and inference services into production systems. Collaborate with cross-...Remote jobMonday to FridayShift work
$130.6k - $192k
About the TeamCome help us build and develop tools serving hundreds of engineers internally! We’re looking for a Fullstack Software Engineer to join our Developer Insights team. About the RoleOur mission is to improve the developer experience of engineers at DoorDash by...Hourly payWork at officeLocal areaRemote workFlexible hours$38.45 - $62.5 per hour
Software Development Engineer in Test (SDET) Location: Remote or Hybrid in TX (Dallas metro-area) Long-Term Contract (3-4 years) Pay Rate: $3... ...interviewing at ConsultNet Technology Services and Solutions by 2x Inferred from the description for this job Medical insurance...Long term contractContract workWork experience placementRemote work- Software Engineer Chalk is building the data platform that powers the future of machine learning applications. We tear down complexity, latency... ...programs in order to optimize arbitrary user Python code, infers and orchestrates infrastructure implied by the structure of that...Work at officeFlexible hours
$80k - $107k
GPU Software Engineer (CUDA) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud... ...maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...Flexible hours
- ...GitHub Actions), K8s, managed Postgres, and AI inference. Today, we serve 500+ customers on our managed cloud. We're now hiring engineers to help us with our Postgres, GitHub... ...tools / techniques more and more during our software development processes. We'd like to share...Immediate start
- Software Engineer II Technology is at the heart of Disney's past, present, and future. Disney Entertainment and ESPN Product & Technology... ...time-series data, model training on GPU clusters, real-time inference pipelines, and model improvement. You will partner with engineering...
- ...how value compounds across the platform. Engineers here use AI as a force multiplier in... ...moments of their lives. About the Role As a Software Engineer, you'll work directly with our... ...that supports production-level inference, evaluation, and monitoring Work across...Temporary workWork at officeFlexible hours
$75.5 - $102 per hour
...make a profound impact, empowering every engineering team with safe, governed, and cost-... ...-7 years of professional experience in software development. Strong proficiency in Python... ...model routing, semantic caching, and batch inference. Experience building developer...Hourly payTemporary workFlexible hours- ...an exciting and rewarding opportunity for you to take your software engineering career to the next level. The Chief Data & Analytics Office... ...secure and high-quality production code for machine learning inference and training systemsProduces architecture and design...Work at office
$160k - $240k
Senior Software Engineer - RDF Infrastructure Location New York Business Area Engineering and CTO Ref # 10054258 Description... ...isolation semantics, statistics and cardinality estimation, inference, replication, scalability, and operational reliability.The...Temporary workFor contractorsWork experience placement$132.6k - $192.3k
...for this position. Skills and Competencies ~5+ years of software engineering experience designing, coding, testing, and operating... ...services, application programming interfaces, data pipelines, and inference pipelines that support real-time and batch AI workloads...Full time- ...Nscale is seeking a Principal AI Engineer to lead the inference and post-training pillar of our AI systems engineering organization in the United States. You will shape the multi-year technical roadmap for serving models and post-training workloads, guiding 20–50+ engineers...
$225k - $300k
...training data to Fortune 500s like Siemens extracting decades of engineering records, Datalab is where businesses turn to when extraction... ...document-understanding systems. That includes building core inference workflows, creating intuitive UI for complex parsing tasks, and...Local area$91.7k - $163.7k
Sr Software Engineer Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives... ..., metadata enrichment, and performance tuning for scalable inference Partner with broader analytics and AI teams to make recommendations...Minimum wageFull timeWork experience placementLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!
- software engineer internship remote New York, NY
- senior software engineer ruby on rails New York, NY
- software developer positions New York, NY
- intermediate software engineer New York, NY
- agile software developer New York, NY
- software engineer intern New York, NY
- part time software developer New York, NY
- rust software engineer New York, NY
- software engineer no experience New York, NY
- software engineer internship New York, NY



