AI Inference Engineer
$230k - $350kPremier Global Links LLC
About the Role
Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.
This role is focused on LLM inference, model serving, distributed systems, and inference runtime performance . The ideal candidate has hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued.
Key Responsibilities
- Build and optimize production LLM inference and model-serving systems .
- Develop inference runtime components using Rust and other systems-level technologies.
- Design and implement batching, scheduling, request routing, and serving infrastructure.
- Build and optimize KV cache and prefix caching systems.
- Scale inference workloads across multi-GPU and multi-node environments.
- Profile, benchmark, and optimize latency, throughput, reliability, and cost.
- Work with inference engines such as vLLM, SGLang, or TensorRT-LLM .
- Investigate performance bottlenecks across the inference stack.
- Contribute to core architecture and technical decisions for the platform.
- Collaborate with a small, hands-on engineering team in a fast-paced environment.
Required Qualifications
- 2–10 years of experience in backend, distributed systems, or systems engineering.
- Hands-on experience building, operating, or optimizing LLM inference or serving systems .
- Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling.
- Experience with a production inference engine such as vLLM, SGLang, or TensorRT-LLM .
- Strong programming experience with Rust, C++, Go, or systems-level Python/PyTorch .
- Experience building performance-critical systems where latency, throughput, and cost are important.
- Strong distributed systems and production software engineering fundamentals.
- Ability to work on-site 5 days per week in Palo Alto, CA.
Preferred Qualifications
- Production Rust experience.
- CUDA or Triton kernel development experience.
- Multi-GPU or multi-node serving experience.
- Experience with NCCL, NVLink, or RDMA .
- Experience with prefix caching, speculative decoding, or prefill/decode disaggregation.
- Contributions to open-source inference projects such as vLLM, SGLang, or Dynamo .
- Experience working on inference systems at an AI provider, accelerator company, research lab, or similar organization.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field.
Technology Environment
Rust | C++ | Go | Python | PyTorch | vLLM | SGLang | TensorRT-LLM | CUDA | Triton | NCCL | NVLink | RDMA | LLM Inference | Distributed Systems
Compensation & Benefits
- $230,000–$350,000 annual salary , based on experience and qualifications.
- Equity opportunity starting at approximately 0.5% , with flexibility based on experience.
- Professional growth and development opportunities.
- High-impact work within a fast-paced AI technology environment.
Work Arrangement
On-Site – Palo Alto, CA
Employees are expected to work from the Palo Alto office 5 days per week .
Equal Opportunity Employer
Premier Global Links LLC is an equal opportunity employer. Qualified applicants are considered based on their skills, experience, education, and qualifications.
- ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software team is the central nervous system of SpaceX - we create mission critical applications...SuggestedPermanent employmentTemporary workRemote workWorldwideWeekend work
$117.7k - $221.4k
...practical, and cost efficient for embodied AI systems. We believe the next generation... .... This operating model reflects how Cola engineers think: build durable intermediate artifacts... ...the data processing, featurization, and inference foundations that power scalable world...SuggestedFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient... ...CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution... ...latency and maximize memory bandwidth on AI accelerators. Write production-level,...SuggestedTemporary workRelocation package
- ...Machine Learning Engineer LiveX AI is building the next generation of realtime, interactive AI avatars—lifelike digital humans that see... ...of diffusion, video, and multimodal models, to low-latency inference optimization for live, streaming deployments. Your work...Suggested
$274k - $300k
...Job Description Job Description Saviynt's AI-powered identity platform manages and governs human and non-human access... .... For more information, please visit AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs...Suggested$120k - $220k
...news and information powered by advanced AI, recommendation systems, and adtech.Recognized... ...sharper every cycle.We're hiring the engineer who owns this agent end-to-end.What You’... ...debugged them, trained a LoRA, optimized inference, built a ComfyUI workflow you’d defend in...Full timeLocal areaWork from home$144k - $236k
...business needs of the team.Responsibilities: AI is at the core of how LinkedIn connects... ...trust platforms. As a Senior AI Software Engineer you will own end-to-end machine learning... ...efficiency or quality improvement (i.e. inference/training efficiency, engineer velocity, or...For contractorsWork at officeImmediate startFlexible hours$171k
Sr. Staff ll, AI Engineer - Search (L7-2)We exist to wow our customers. We know we’re doing the right thing when we hear our customers say... ...infrastructure, including vector databases, high-throughput inference, and real-time data pipelines for context enrichment and low-latency...Temporary workFlexible hours$175k - $287k
...business needs of the team.Responsibilities: AI is at the core of how LinkedIn connects... ...trust platforms. As a Staff AI Software Engineer you will own end-to-end machine learning... ...efficiency or quality improvement (i.e. inference/training efficiency, engineer velocity, or...For contractorsWork at officeImmediate startFlexible hours$170k - $216k
...speed up developer velocity. We’re looking for a software engineer to join the team to build and maintain the critical data and... ...Staff Software Engineer. You will: Develop Waymo's inference platform to make it scalable, high throughput, and low latency...Full timeRemote work- ...the forefront of a new era in enterprise AI — one defined not by model capability... ...of frontier AI research and production engineering — investigating the foundational challenges... ...persistence architectures, model selection and inference routing strategies, autonomy and goal-...Full timeWork experience placementLive inWork at officeLocal areaRelocation
$286.4k - $358k
Uniphore is the Business AI company. Our sovereign, composable and secure AI platform... ...Business AI. We are seeking a VP of AI Engineering to lead the architecture and delivery of... ...that support SLLM fine-tuning, inference, prompt engineering, and RAG pipelines in...Full time- ...opportunity for you to take your software engineering career to the next level. As a Software... ...systemsLeverages enterprise-authorized AI coding assist tools within the work environment... ...TorchServe, TensorFlow Serving, Triton Inference Server)Familiarity with distributed...
- ...come to the right place. As a Principal Software Engineer at JPMorganChase within the Corporate Sector – AI/ML & Data Platforms for LLM Suite, you will lead a... ...detection, versioning, and rollback Optimize inference for latency, throughput, caching, batching, model...
- ...is a research lab of top researchers and engineers, building the world’s top-ranked realtime... ...used to power the largest consumer-facing AI applications available, across categories... ...-of-the-art models, optimizing realtime inference, and creating best-in-class APIs and products...Full timeContract workWork at officeRelocation
$142.8k - $274.8k
...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s... ...Systems organization is developing AI-native silicon and hyperscale systems designed... ...enable industry-leading AI training and inference. The Platform Systems Engineering (PSE)...Ongoing contractWork at officeLocal areaWorldwide3 days per week$151.3k - $283.8k
...technological advancements such as cloud, AI, and network security. While driving the... ...context of Large Language Model (LLM) inference and training.2.Operator & Performance Optimization... ...: Master’s or Ph.D. degree in Computer Engineering, Electronic Engineering,...Full timeRelocation package$207k - $300k
...design for complex 1-6-month Forward Deployed Engineer (FDE) embeds and 2-4-week Strike Sprints... ...).Experience integrating generative AI tools or LLM interfaces into workflows.Preferred... ...model efficiency pipelines, optimize inference serving engines, and establish graceful...- ...Description Gauss Labs builds Industrial AI for the world's leading manufacturers,... ...production systems. As a Senior/Staff AI Engineer, you will turn ML research into robust, scalable... ..., and production: data, training, and inference pipelines, CI/CD, observability, and...Full timeShift work
$200k - $270k
...Job Description Samsung SDS America AI Team is researching the next generation of... ...We are looking for a Senior Physical AI Engineer to join the team developing this end-to-end... ...similar platforms Optimize low-latency inference and control systems for real-world robotic...WorldwideFlexible hours$160.36k - $240.54k
...driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets,... ...data generation to on-road validation.Maintain an in-house ML inference platform to serve large language models efficiently.Maintain an...Immediate startFlexible hours$207k - $300k
Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically... ...qualifications:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related...$100k
...is leading the industry on cutting-edge AI technology, revolutionizing performance expectations... ...Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth... ...technologies for next-generation AI inference and training clusters. This role is on-site...Permanent employment- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and... ...advance your career. THE ROLEWe are hiring AI Engineers to build recursive self-improvement... ...JAX, TensorFlow, or distributed training/inference systems.Experience with reinforcement learning...
- ...Description Job Description Accellor is an AI-native services firm purpose-built for... ...outcomes through advanced AI, data, and engineering capabilities. Our mission is to... ...for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale...
- ...future of mission critical workflows by leveraging the latest in AI. Our purpose is to solve humanity's hardest hurdles, starting... ...select angels. The Opportunity You'll be part of the core engineering team building scalable, high-impact systems that power our healthcare...Work experience placementInternship
- ...Embodied AI Engineer UnitX builds the world's leading physical AI systems to automate repetitive visual tasks in factories. UnitX is... ...deploying, profiling, and optimizing ML models for real-time inference on robotic hardware (e.g., NVIDIA Jetson, TensorRT, CUDA). ~...
$180k - $275k
...About the role Own verification of our AI compute core - tensor pipelines, MAC arrays, accumulator logic, and the compute ↔ memory interconnect. Work with chip-design and software teams driving DensityAI's AI accelerator program from first silicon through scale-out...H1bVisa sponsorshipWork visa- ...Job Description We’re hiring an Explainable AI Engineer to build, verify, and validate business impact predictions based on outputs of our AI Engines. You’ll have an opportunity to learn from Stanford and Georgia Tech professors and work across algorithms and data...Full timePart timeFor contractorsRemote work
$2,000 per month
...Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale - unleashing... ...Agentic Workflows. We are looking for an innovative Agentic AI Engineer to join our team to build autonomous, enterprise-grounded...Local areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Inference Engineer. Be the first to apply!


