Tech Lead - AI Inference
WEKA
Role Description
We are seeking a Tech Lead to lead our AI Inference team. In this role, you will bridge the gap between complex research and production-grade engineering, while cultivating a high-performing team culture. You will lead and grow a squad of 3 developers, balancing hands-on technical contribution with strong people leadership — setting direction, unblocking your team, and driving execution on high-performance systems that optimize Large Language Model (LLM) serving.
The ideal candidate combines deep technical expertise in inference and scale with the leadership maturity to mentor, motivate, and develop engineers in the evolving ecosystem of serving frameworks like vLLM and LMCache.
What You'll Work On
- Lead & Own: Take end-to-end ownership of AMG's core inference infrastructure — from the NVMe Token Warehouse and GDS data paths to the vLLM/LMCache serving stack — driving technical decisions and delivery outcomes.
- Technical Direction: Guide a team of engineers through design, implementation, and delivery of high-throughput, low-latency LLM inference systems, setting high standards for code quality, architecture, and reliability.
- Build at Scale: Stay hands-on across the AMG stack (Python, C++, CUDA, vLLM, NIXL/Dynamo, Kubernetes), contributing directly to production systems while providing technical leadership to the team.
- Solve Hard Problems: Tackle the real frontier challenges of inference engineering — disaggregated prefill/decode, persistent off-HBM KV caching, RDMA-based transport, and multi-tier GPU memory hierarchies — that define what's possible at scale.
- Grow People & Teams: Mentor and coach engineers through regular 1:1s, career coaching, and sprint reviews. Foster a culture of ownership, collaboration, and technical excellence within the AMG team.
- Stay on the Frontier: Track the evolving inference ecosystem, benchmark new tools (SGLang, TRT-LLM, NVIDIA Dynamo), and help the team make timely decisions about when to adopt, build, or pivot.
Qualifications
- 5+ years of professional software engineering, with proven experience leading engineers and owning complex production systems — ideally in AI/ML infrastructure or high-performance computing.
- Hands-on expertise with LLM serving systems — KV cache reuse, disaggregated prefill/decode, continuous batching, and multi-tier GPU memory hierarchies (HBM → NVMe). Strong familiarity with vLLM, LMCache, NIXL/NVIDIA Dynamo, or similar frameworks.
- Strong Python and C++ skills (Rust a plus), with a solid grasp of CUDA, GPU memory management, and high-performance I/O — including GPUDirect Storage (GDS), RDMA, and NVMe data paths.
- Experience deploying and scaling GPU workloads on Kubernetes, with familiarity in RDMA networking, bare-metal GPU clusters (H100/A100), and high-throughput distributed storage.
- Demonstrated ability to mentor and develop engineers — running effective 1:1s, supporting career growth, and balancing technical execution with long-term team health.
- A strong sense of engineering craftsmanship, with a track record of building reliable, high-throughput systems and continuously improving engineering practices.
Benefits
- Medical, Dental, Vision, Life Insurance
- 401(K)
- Flexible Time off (FTO)
- Sick time
- Leave of absence as per the FMLA and other relevant leave laws
Company Description
WEKA is an equal opportunity employer that prohibits discrimination and harassment of any kind. We provide equal opportunities to all employees and applicants for employment without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.
$2,000 per month
...About Etched Etched is building AI chips that are hard-coded for individual model architectures. Our first product (Sohu) only... ...scheduling logic for handling continuous batching and real time inference Implement inference-time acceleration techniques such as speculative...SuggestedFull timeWork at officeRelocation package$92k - $135k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... ...Our technology provides enterprises and leading AI labs with the most performant, efficient... .... What You’ll Do: Join the Inference team to ship production features that improve...SuggestedPermanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$152k - $287.5k
...seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design... ...GPU-accelerated software that powers today’s most sophisticated AI applications. Our team is responsible for developing and maintaining...SuggestedFull time$320k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and... ...AI systems. About the Role Our mandate is to make inference deployment boring and unattended. Anthropic serves Claude to...SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work- ...frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic... ...), especially how they influence latency and throughput of inference. ~ Strong understanding or working experience with distributed...SuggestedFull timeWork experience placementWork at officeRemote workFlexible hours
- Role Description We are looking for an AI Tech Lead to lead the design, engineering, and rollout of AI-powered solutions, with a strong... ...unsupervised learning, model training, feature engineering, evaluation, inference, and model performance metrics. ~Experience with MLOps /...Full timeCasual workRemote work
- ...optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and endpoint devices,... ...code and conventional C++ DSP and control code. Role: The AI Inference Engineer in Quadric is the key bridge between the...Full timeTemporary workWork from home
- ...seeks an Engineering Technical Lead to join our team. This role is... ...the next generation of AI CX tools — automated QA, conversational... ...real-time platform. ~A tech lead here is a player-coach —... .../ML pipelines (RAG, streaming inference, eval) or connectors (CCaaS/CTI...Full timeRemote work
$139.2k - $174k
Role Description DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key technical leader responsible for...Full timeRemote work- ...A leading technology company in California is seeking an AI Inference Engineer to bridge AI models with unique platforms. Key responsibilities include model optimization, deployment, and performance profiling. Candidates should have a Bachelor’s or Master’s degree, 5+...Full timeRemote workWork from home
$170k - $200k
Role Description As a Full-Stack Rails Tech Lead, you’ll architect and implement LLM-powered... ...combines expert Ruby on Rails skills with AI integration expertise, requiring you to... ...AI-driven interactions. ~Optimize LLM inference for low latency and cost, using caching and...Full timeWork at officeRemote work$221k - $260k
...Staff Machine Learning Engineer – AI Tech Lead Location: USA The proliferation of AI and machine log data has the potential to... ...Design scalable LLMOps and AI agent infrastructure, including inference routing, latency optimization, cost control, and production...Full timeWorldwide$320k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and... ...build beneficial AI systems. About the role The Cloud Inference team scales and optimizes Claude to serve the massive audiences...Full timeWork at officeVisa sponsorshipFlexible hours- ...is seeking a Principal Software Engineer, Tech Lead - Platform to join the team, reporting to... ...consumption increasingly shifts toward AI and agentic applications, you will also shape... ..., such as training pipelines or inference for LLM-powered and agentic tools. Benefits...Full timeFlexible hoursShift work
- ...era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the... ...learning systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing...Full time
$160k - $200k
...Mitre Media is redefining FinTech with AI-driven tools that empower millions of investors... ...About the Role As a Full-Stack Rails Tech Lead, you’ll architect and implement LLM-... ...AI-driven interactions. Optimize LLM inference for low latency and cost, using caching and...Work at officeRemote work$200k - $250k
...About Wizard AI At Wizard AI, we’re building a high-performing AI Shopping Agent that helps people discover the best products... ...to monitoring, performance, and scaling — across a custom-built inference platform powering a live conversational product. This isn’t a...Full timeRemote workFlexible hours- ...the future of embodied intelligence. Our AI-driven systems enable robots to adapt, learn... ...a Senior Machine Learning Engineer (Tech Lead) for this new team. You write code, set the... ...project. Strongly Preferred: Edge inference depth (TensorRT, ONNX, edge-class...Full timeFlexible hours
$114.8k - $191.4k
...leader who thrives in an agile environment, embraces the future of AI-driven development, and enjoys mentoring a high performing team... ...code analysis, and problem-solving. ~Serves as the technical lead, owning the development of the entire team — from junior to senior...Permanent employmentFull timeRemote workFlexible hours$96.95k - $130k
...service of mission partners across the globe. Mission Technologies is leading the next evolution of national defense - the data evolution - by... ...and commercial customers. Our capabilities range from C5ISR, AI and Big Data, cyber operations and synthetic training...Full timeContract workWork at officeLocal areaWorldwide$180k - $200k
Role Description As a Software Engineering Manager/Tech Lead at Cylinder, you will be a technical leader and owner of feature delivery, helping... ...part of this role involves leading your team in building AI-powered features, collaborating with our data and clinical teams...Full timeRemote work$115k - $183k
Role Description We are seeking a highly skilled Senior AI Tech Lead, this is a role requiring DevOps Engineering experience with a deep... ...Python, PySpark, and Kafka for both batch and real-time inference. ~Leverage advanced orchestration frameworks such as LangChain...Full time- Role Description We are seeking a Technical Lead who is equally at home writing production code and leading a team. This is a hands-on... ...team’s engineering process, mentoring engineers, and championing AI adoption. You will bridge technical execution and business strategy...Full timeRemote work
$200k - $230k
Role Description The Tech Lead is the engineering anchor for both Fool.com and Motley Fool Money, responsible for setting technical direction... ...to translate business goals into technical strategy. ~Drive AI integration thoughtfully, both in production features and in how...Full timeRemote workFlexible hours- ...live OutSystems products and SaaS, while leading the delivery of new capabilities that solve... ...least 1 year leading technical delivery as a Tech Lead (required) ~Strong experience in... ...performance (required) ~Active use of AI-augmented development in real work — Claude...Full time
- ...operational support to Video Template Review teams. Owns tooling, AI-assisted workflows, technical standards, and review processes... ...Pro templates, DaVinci Resolve templates, and related assets. ~Lead technical upskilling and tooling adoption for Video Template Reviewers...Full timeRemote workWork from homeHome officeFlexible hours
$227.84k - $335k
...innovative solutions as part of Twilio’s incubation team. As a Tech Lead, you’ll work closely with cross-functional partners to design... ...prototypes and robust, production-ready solutions. ~Spearhead AI R&D: Drive the technical direction and collaborate with cross-functional...Full timeLocal areaRemote workShift work- Role Description The Tech Lead — AI Applications is the senior technical leader for the product-suite implementation tier of an AI-native retail decisioning platform — the commercial layer that delivers measurable business outcomes through a portfolio of product suites...Full timeContract workRemote work
- Role Description As a Tech Lead, you will play a strategic role, serving as a technical reference and team leader. This position requires... ...support in solving complex problems. ~Promote the adoption of AI-assisted engineering practices within your squad and chapter....Full timeRemote workHome officeFlexible hours
- Role Description ~Write clean, production-grade Python across AI integrations, backend services, and RESTful APIs. ~Implement and... ...calls, technical proposals, scoping, and client-facing demos. ~Lead architecture reviews, produce technical design documents, and contribute...Full timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tech Lead - AI Inference. Be the first to apply!


















