Staff Software Engineer, AI Inference
$190k - $230kTLA Tech
About Syllo
Syllo is on a mission to transform litigation. Our product is a unified litigation platform that enables lawyers and paralegals to safely harness the power of language models and agentic AI throughout the litigation life cycle. Since going to market, we have gained a diverse group of enterprise customers, including some of the biggest law firms and corporations in the country, and we are quickly expanding. By reducing the expense of litigation industry-wide, we aim to improve access to high-quality representation and promote the alignment of legal outcomes with merit.
About the Role
As we continue to scale our AI platform, we're investing in our own inference stack to deliver best-in-class performance, reliability, cost efficiency, and flexibility across the latest generation of open-source language models.
We're looking for a Staff Software Engineer to spearhead this effort.
You'll define the architecture, evaluate emerging technologies, and build the systems that power model serving. You'll partner closely with machine learning, infrastructure, and product engineering to establish the foundation for how AI models are deployed, optimized, monitored, and operated in production.
This is a highly hands-on technical role. You'll spend the majority of your time designing, building, and optimizing production systems while helping shape our long-term AI infrastructure strategy.
Responsibilities
- Lead the design and development of our production inference platform.
- Define the technical roadmap for inference infrastructure, model serving, and runtime optimization.
- Build and operate scalable, cost-effective systems for serving large language models in production.
- Evaluate and integrate modern inference technologies, frameworks, and serving runtimes.
- Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost.
- Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence.
- Partner with ML engineers to productionize new models and inference techniques.
- Establish benchmarking methodologies to evaluate new models, runtimes, and hardware.
- Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions.
- Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.
Qualifications
- Significant experience designing and operating production AI inference systems.
- Experience building or leading production LLM serving infrastructure.
- Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies.
- Strong background in distributed systems, backend infrastructure, or high-performance platform engineering.
- Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency.
- Experience operating GPU infrastructure in production.
- Strong proficiency in Python and at least one systems programming language (such as Go, Rust, or C++).
- Proven ability to lead technical architecture for complex infrastructure initiatives.
- Excellent communication skills and the ability to influence technical direction across engineering teams.
Salary Range ($190- $230K) plus health insurance and equity.
United States - Remote Pay Range
$190,000—$230,000 USD
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and... ...and unattended. As a Software Engineer on the Launch... ...Currently, we expect all staff to be in one of our...SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work- ...enterprises who are building AI systems to power magical experiences... ...is a team of researchers, engineers, designers, and more, who are... ...looking for Members of Technical Staff to join the Model Serving team... ...latency and throughput of inference. ~ Strong understanding or working...SuggestedFull timeWork experience placementWork at officeRemote workFlexible hours
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...role The Cloud Inference team scales and optimizes... ...Have significant software engineering experience,... ...Currently, we expect all staff to be in one of our offices...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours- ...of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the... ...systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between...SuggestedFull time
$167.2k - $209k
...world. DigitalOcean is expanding its AI Infrastructure layer to support the... .... We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you... ...to be able to develop high quality software while availing of all the productivity...SuggestedFull timeLocal areaRemote workWorldwideFlexible hours$92k - $135k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... .... What You’ll Do: Join the Inference team to ship production features that improve... ...quickly with mentorship from experienced engineers. About the role: Implement well...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$2,000 per month
...Etched is building AI chips that are hard-coded... ...design of the Sohu host software stack Implement high... ...batching and real time inference Implement inference-... ...Cupertino, and greatly value engineering skills. We do not have... ...all of our technical staff to contribute to both...Full timeWork at officeRelocation package$216k - $258k
...full potential of their data for AI applications. The platform... ...infrastructure that power AI inference at 100K+ QPS with millisecond... ...Evolve Tecton’s query execution engine to support complex, multi-stage... ...distributed and/or highly concurrent software systems ~ Degree in Computer...Full timeRemote workFlexible hours- ...entertainment company and the leader in AI music. We are backed by... ...team at the Senior and Staff levels. These aren’t maintenance... ...features end-to-end, drive engineering quality, and contribute to architectural... ...features, and real-time AI inference all live on this platform. If...Full timeWork at officeLocal areaWorldwide
$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference...Temporary workFor contractorsWork experience placement- ...foundation of a neurosymbolic AI agent that learns more about... ...Join Onton as a Founding Engineer and set the strategic foundation... ...out a performant and scalable inference engine to support more... ...and is passionate about making software tools accessible to all, we want...Full timeWork at officeLocal areaRemote workRelocation3 days per week
$152k - $241.5k
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize... ...software that powers today’s most sophisticated AI applications. Our team is responsible for developing...Full timeRemote work$238k - $302k
...metrics development teams including large AI deployments, System engineers and Data scientists that produce... ...not limited to: human triage, VLM inference, clustering. Key workstreams that the... ...have: ~7+ years of full-time software engineering experience, or a...Full timeRemote work$229k - $343k
...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store... ...features reliably for large-scale batch inference and low-latency online inference, maintaining... ...with responsible use of emerging AI technologies to improve engineering...Full timeLive inWork at officeLocal area- ...Description:DataRobot delivers AI that maximizes impact and... ...DataRobot’s Fleet team is the engine behind how our platform runs... ...That’s where you come in.As a Staff Software Engineer, you’ll be responsible... ...for training and inference.Why Join the Fleet Management...Full timeLocal areaRemote workWorldwideFlexible hours
- ...00.00 - $249,500.00Job DescriptionThe Staff Software Engineer is a position of technical expertise,... ...through production implementation.LLM & AI Platform Architecture: Define how LLMs... ...of LLM architectures, model serving, inference, evaluation, orchestration, and integration...Full timeFlexible hours
$186k - $280k
...innovate, grow, and thrive with our open, AI-driven commerce ecosystem. As the... ...commerce, this is the place for you. The Staff Software Engineer is a senior technical leader within... ...databases, model orchestration tools, inference frameworks, and cloud-native ML...Full timeWork at officeLocal areaRemote workWorldwide3 days per week$163k - $253k
...RoleWe are seeking a Senior Staff Engineer to build and optimize the EDA... ...processes for EDA licensing and R&D software. In this role, you will... ...and cost efficiency at scale.AI/PI Group is an internal AI and... ..., routing strategies, and inference configurations that meet product...Flexible hours$242k - $389k
..., WA / Remote (United States)Software – Software Systems /Full-time... ...alongside a team of strong software engineers and act as a force multiplier... ...cutting-edge ML Training OR Inference performance optimization... ...use artificial intelligence (AI) tools to support parts of...Full timeRemote work$349k
...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...the future. We are seeking a Staff Engineer to help our development of... ...of AI training and inference at scale.As a Staff Engineer... ...Qualifications10+ years of experience in software engineering, platform...Work at officeLocal areaImmediate startWork from homeFlexible hours- ...reliable on-demand, logistics engine for last-mile retail... ...to join our team. As a Staff Machine Learning... .... Proficiency in using AI coding tools (e.g., Claude... ...Codex, Cursor) in the full software development lifecycle,... ...applied ML for Causal Inference and Recommendation...Hourly payWork at officeLocal areaRemote workFlexible hours
$193.3k - $261.5k
...builds Amazon Neuron, the software development kit used to... ...unparalleled ML inference and training performance... ...software boundary, our engineers build systematic infrastructure... ...of what's possible in AI acceleration.As part of... ..., supervisors, and staff; adhere to standards of...Work experience placementInternshipLocal areaFlexible hours$198k - $326k
...responsible for scaling LinkedIn's AI model training, feature engineering and serving with hundreds... ..., data infra, compute software, and hardware to harness... ...cloud, enable GPU based inference for a large variety of... ...at scale.As a Sr. Staff Software Engineer, you will...For contractorsWork at officeFlexible hours- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...groundbreaking AI training and inference possible.The Lambda Infrastructure Engineering organization forges the... ...Role:We are seeking a seasoned Staff Storage Software Engineer with deep experience designing...Work at officeLocal areaWork from homeFlexible hours
$151.8k - $332.2k
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on...Full timeWork at officeRemote work$188k - $275k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave... ...Do Description of the team: The Inference team is responsible for delivering high-performance... ...: We are looking for an Applied AI Engineer to help us understand, measure, and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$265k
.... And with JunieAI, our purpose-built AI platform, we’re reimagining how private... ...some or all of the time. Senior Staff Software Engineer Location: Americas (preferably EST) —... ...support production AI workloads, from inference pipelines to agent fleets Design systems...Full timeWork at officeLocal areaRemote workWork from homeFlexible hours$190k - $280k
...of any specified location above. We are AI Native We are building an AI native company... ...sure the seams don't show. You own the inference pipeline end to end: model selection, serving... ...don't run away from us. You are a cloud engineer who is deeply fluent in how AI systems...Work at officeRemote workFlexible hours$184k - $230k
...Business Area: Engineering Seniority Level: Mid-Senior level... ...- enabling data and AI workloads to run anywhere, without... ...belong. We are seeking a Staff Software Engineer to lead the... ...: Lead the deployment of inference servers (vLLM, Triton) using...Work from homeFlexible hours$200k - $350k
...across multiple systems and platforms. We are seeking a Staff AI Engineer to lead the design and deployment of production AI systems... ...architecture for production AI systems. Design model serving, inference, evaluation, and optimization pipelines. Improve model...Full timeImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference. Be the first to apply!


