Software Engineer - Training/Inference (C++)
$180kSpaceXAI
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. About the Role SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability. As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency). This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale Responsibilities: Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache). Optimize latency and throughput of model inference under real production workloads. Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency. Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation). Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels. Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates. Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.
BASIC QUALIFICATIONS
Deep low-level systems programming (C/C++ or Rust) Experience with large-scale, high-concurrent production serving. Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.). Strong background in system optimizations: batching, caching, load balancing, parallelism. Low-level inference optimizations: GPU kernels, code generation. Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics. Experience with testing, benchmarking, and reliability of inference services. Experience designing and implementing CI/CD infrastructure for inference.COMPENSATION AND BENEFITS
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice. #J-18808-Ljbffr SpaceXAI$143.7k - $194.4k
...Neuron is the complete software stack for AWS... ...the Machine Learning Inference Applications team to build... ...on Neuron chips.As an engineer on this team, you'll work... ...and services using: C#, C++, Java, or Perl experience... ...architecture, training/inference lifecycles,...TrainingC++InternshipFlexible hours- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to ship AI... ....2 Live draft model training for speculative decoding... ...languages, such as Python or C++. Familiarity with... ...the performance of software systems, particularly...TrainingC++Flexible hours
- ...thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the... ...and orchestration. Ensure software solutions are optimized for... ...programming languages such as Python, C++, Rust, or Golang... ...compensation, promotion, benefits, training, discipline, and termination...TrainingC++Full timeLocal areaImmediate start
- ...retrieval and ranking to multi-agent LLM engines, post-training infrastructure, and personalized... ...signals.- Contribute to Training and Inference Infrastructure: Collaborate on post-training... ...online services.- Proficiency in C++, Go, or Java (C++ preferred); strong systems...TrainingC++
$168.1k - $227.4k
...of hardware, firmware, and software to deliver unparalleled virtualization... ...for high-performance training and inference workloads.We are looking for an experienced software engineer to drive development for new... ..., and expertise in C/C++ or Rust development in a Linux...TrainingC++InternshipFlexible hours$143.7k - $194.4k
...talented scientists and engineers to innovate on behalf... ...stable and efficient training system for model training... ...professional software development experience... ...architecture, training/inference lifecycles, and optimization... ...Experience with CUDA/C++/Kernel developmentAmazon...TrainingC++InternshipFlexible hours$143.7k - $194.4k
...scale systems, solving critical engineering problems, and delivering... ...non-internship professional software development experience- 2+ years... ..., and services using: C#, C++, Java, or Perl experience- 1... ..., including architecture, training/inference lifecycles, and optimization...TrainingC++InternshipFlexible hours$168.1k - $227.4k
...hardware, firmware, application software and services to deliver new... ..., and expertise in C, C++ or Rust development in a Linux... ...Collaborate with hardware engineering teams to influence future... ...to power high-performance training and inference workloads. Basic qualifications...TrainingC++InternshipFlexible hours$168.1k - $227.4k
AWS Neuron is the complete software stack for the AWS Inferentia... ...the Sr. Software Development Engineer for the Neuron Foundation Tools... ...to ensure that the our C++ compiler and runtime generates... ...build massive-scale distributed training and inference solutions. This organization...TrainingC++InternshipWork from homeFlexible hours$143.7k - $194.4k
...Amazon builds AWS Neuron, the software development kit used to... ...JAX enabling unparalleled ML inference and training performance.The Inference Enablement... ...-software boundary, our engineers build systematic... ...Software development experience in C++, Python (experience in at least...TrainingC++Work experience placementInternshipFlexible hours- ...solutions, including large-scale model training, high-performance inference, feature platforms, and embedding... ...We are looking for exceptional Software Engineers who are passionate about AI infrastructure... .... Strong programming skills in C++, Python, Java, Go, or similar...TrainingC++Internship
$143.7k - $194.4k
...used across 23 countries. The engineering model is changing fast: we... ...tools as your default mode of software delivery.Build and maintain... ..., and services using: C#, C++, Java, or Perl experience- Bachelor... ..., including architecture, training/inference lifecycles, and optimization...TrainingC++InternshipFlexible hours$143.7k - $194.4k
...selection. This is AI-first engineering: you'll build the... ...design through training, evaluation, and production... ...internship professional software development experience... ...services using: C#, C++, Java, or Perl experience... ..., training/inference lifecycles, and optimization...TrainingC++InternshipFlexible hours$168.1k - $227.4k
...build Amazon Neuron, the software development kit used... ....As a Senior Software Engineer on our Machine... ...building distributed inference support for Pytorch in... ...Knowledge of Python and/or C++ programming- 5+ years... ...transformer architecture, training/inference lifecycles,...TrainingC++InternshipWork from homeFlexible hours$165.2k - $223.6k
...alongside talented scientists, engineers, and technical program... ...looking for an experienced Software Development Engineer with expertise... ..., and services using: C#, C++, Java, or Perl experience- 1... ...transformer architecture, training/inference lifecycles, and optimization...TrainingC++InternshipLocal areaFlexible hours$143.7k - $194.4k
...) is building a central pipeline of Software Development Engineer (SDE) talent for anticipated roles in... ..., systems, and services using: C#, C++, Java, or Perl experience- Bachelor'... ...including transformer architecture, training/inference lifecycles, and optimization techniques...TrainingC++InternshipFlexible hoursDay shift$158.1k - $213.8k
...responsibilitiesWe're looking for a Software Development Engineer to join the Analytics & Insights team... ..., systems, and services using: C#, C++, Java, or Perl experience- 1+ years... ...including transformer architecture, training/inference lifecycles, and optimization techniquesAmazon...TrainingC++InternshipWorldwideFlexible hours$143.7k - $194.4k
...enduring challenges in the industry.As a Software Development Engineer in the AWS Healthcare AI team, you'... ..., systems, and services using: C#, C++, Java, or Perl experience- 1+ years... ...including transformer architecture, training/inference lifecycles, and optimization...TrainingC++InternshipWorldwideFlexible hours$143.7k - $194.4k
...help.You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists,... ...one modern language such as Java, C++, or C# including object-oriented design... ...fundamentals, including architecture, training/inference lifecycles, and optimization of...TrainingC++InternshipWorldwideFlexible hours$143.7k - $194.4k
We are looking for a Software Development Engineer II (SDE-2) to join the EKS Runtime Release team.... ...traditional containers to cutting edge AI training and inference.Mentorship: Mentor junior engineers... ...as Python, Ruby, Golang, Java, C++, C#, RustPreferred qualification -...TrainingC++InternshipWork from homeFlexible hours$143.7k - $194.4k
Do you want to build software systems powered by generative AI that... ...a team of passionate engineers and scientists to solve problems... ...systems, and services using: C#, C++, Java, or Perl experience-... ...transformer architecture, training/inference lifecycles, and optimization...TrainingC++Temporary workWorldwideFlexible hours$142.3k - $263.3k
...Seattle, Washington, United States Software and Services We are looking for software engineers to join our small team with... ...languages like Rust and C++. Collaborate with cross-functional... ...storage needs for large-scale training and inference. Implement and manage...TrainingC++Relocation$90k - $180k
...reliable microservices for training/inference, vector indexing, and real-time... ...to design reviews and engineering best practices. Mentor peers... ...primary), plus one of Go/Java/C++ for performance services. Distributed... ...Tech. We’re a team of software engineers, data scientists,...TrainingC++Full timeTemporary workPart time$168.1k - $227.4k
AWS Neuron is the complete software stack for AWS Inferentia and Trainium... ...learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on... ...learning models, their architecture, training and inference lifecycles along...TrainingWork experience placementInternshipLocal areaFlexible hours$143.7k - $194.4k
...the Data Science and Engineering (D:SE) team is the central... ...layer rich enough to train models on, composable... ....We're looking for a Software Development Engineer to... ...Store, and real-time inference endpoints; embedding pipelines... ...services using: C#, C++, Java, or Perl...TrainingC++Contract workWork experience placementInternshipLive inFlexible hours$42.75 per hour
...and operating the end-to-end pipeline —from offline training data orchestration to real-time online inference—that delivers seamless, high-frequency discovery to... ...limited to the following programming languages: C, C++, Java or Golang; Effective communication skills and...TrainingC++Hourly paySummer workInternship- ...Model Optimization & Deployment Engineer, you will focus on bringing... ...and build highly concurrent inference code to ensure real-time, deterministic... ...low latency, and memory-safe C++ and CUDA code for real-time... ...Experience with distributed training pipelines and model/tensor...TrainingC++Temporary workRelocation package
$60 - $65 per hour
Position: Software Engineer IILocation: Seattle, WashingtonDuration: ContractJob ID: 171155Job Overview... ...languages such as Java, Python, C++, or similar.Experience with software development... ...in the market; the skills, education, training, credentials and experience of the...TrainingC++Full time- ...candidates generation, profile generation, training examples generation, realtime online... ...the following programming languages: C, C++, Java or Golang;- Effective communication... ...areas: personalized recommendations, search engine, machine learning, distributed storage system...TrainingC++
$254k - $350k
...Remote (United States)Software - Software Systems /... ...techniques in distributed training, quantization,... ...team of strong software engineers and act as a force... ...-edge ML Training OR Inference performance optimization... ...in Python and C++Experience with model compression...TrainingC++Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - Training/Inference (C++). Be the first to apply!



