AI Inference Platform Engineer
$200k - $250kDRW
DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk.Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters–it's how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.About the Role We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use. You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy. You'll own the serving platform end-to-end: onboarding newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet. What You'll Do Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes. Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems. Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies. Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models. Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing. Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads. Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement. Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments. Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements. Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads. What We're Looking For The Tech Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions. Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang. Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention. Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching. Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes. Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode. Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads. Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana. Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers. Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness. The Intangibles You take a measurement-driven approach to performance optimization. You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries. You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems. You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes. You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance. You communicate clearly and can explain complex performance tradeoffs across engineering teams. The annual base salary range for this position is $200,000 to $250,000 depending on the candidate’s experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus. In addition, DRW offers a comprehensive suite of employee benefits including group medical, pharmacy, dental and vision insurance, 401k (with discretionary employer match), short and long-term disability, life and AD&D insurance, health savings accounts, and flexible spending accounts.For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at . California residents, please review the California Privacy Notice for information about certain legal rights at .[#LI-VD1]
$100k - $140k
...Department BSD CTD - Platform Engineering - GDC About the Department The Center for Translational... ...implementation, security automation, & AI/ML infrastructure management across the... ...machine learning models for inference, optimizing model and hardware performance...SuggestedWork experience placement$165k - $225k
...Moonlite delivers high-performance AI infrastructure for... ...software-defined networking (SDN) platform that enables high-performance... ...distributed computing, model training, inference, and data-intensive workloads... ...– enabling researchers and engineers to access enterprise-grade...SuggestedImmediate startRemote workFlexible hours$165k - $225k
...Moonlite delivers high-performance AI infrastructure for organizations... ...out our GPU-accelerated compute platform that powers distributed AI training and inference, large-scale simulations, and computational... ...-enabling researchers and engineers to programmatically access high-...SuggestedImmediate startRemote workFlexible hours$119k - $169.4k
...While the internal title for this position is Senior Automation Engineer, this role has been posted externally under a different title to... ...artificial intelligence. This role sits at the intersection of advanced AI engineering, security architecture, and organizational strategy...SuggestedFull timeWork at officeImmediate start$124.36k - $146.3k
...excel at—all from Day One.Job DescriptionJob SummaryThe Senior Engineer (Generative AI) is responsible for designing, developing, and deploying... ...practices for production environments3. Cloud, Platform & Scalability EngineeringDevelop and deploy GenAI systems across...SuggestedWork experience placementLocal area3 days per week$137.4k - $233.6k
...technology and exceptional service. Job Title: Cloud Data Platform EngineerPosition Overview:The Cloud Data Platform Engineer is responsible for designing, implementing,... ...supporting enterprise data, analytics, and AI workloads on technologies including Databricks, Snowflake...Full timeH1bWorldwideFlexible hours$176k - $179.5k
...leaders, practice leaders and cutting-edge engineers. Your team is led by some of our firm’s... ...additional features required to deploy this platform for Defense and other Government clients.... ..., knowledge graphs, and/or generative AI will serve you well. You would be working...ApprenticeshipEasy work$147.76k - $240.11k
...connected assets worldwide, our teams use data, technology, advanced analytics, and AI capabilities to help our customers build a better, more sustainable world.Job SummaryThe PIM Platform Engineering Manager leads engineering strategy and delivery for the PIM platform, owning...Full timePart timeWorldwideFlexible hours- ...AI Platform Engineer The Aspen Group (TAG) is one of the largest and most trusted retail healthcare business support organizations in the U.S., supporting over 23,000 healthcare professionals and team members at more than 1,150 locations across 48 states. Our five...Work at officeRelocation
- ...Travel: Occasional international travel required Job Duration: Long Term Contract Job Summary We are seeking an AI Platform Engineer to build and scale enterprise AI platform capabilities that enable multiple product teams. This role focuses on developing production...Long term contractRemote work
- ...Deloitte Tax LLP’s Product Engineering team is leading a transformation to an agentic software development platform. You will own the LLM Wiki and knowledge bases, design scalable AI solutions, and ensure alignment with security and governance standards across AWS Bedrock...
$119.4k - $204.6k
...Essential Responsibilities Define enterprise-wide platform strategy, vision, and target-state architectures... ...standards and guardrails across all platform engineering domains. Drive innovation in cloud, data, and AI platforms (Fabric, OpenAI/Foundry, etc.). Lead...Full timeTemporary workPart timeWork from home3 days per week$145k
...Sydney, Shanghai, Singapore, and London. What you'll do as a Platform Engineer Intern at Akuna: We are seeking Platform Engineer Interns... ..., you may apply to any Trader roles of interest. We are an AI-friendly company and encourage employees to leverage AI in their...Summer workInternshipWork at office- ...Northern Trust is seeking a Sr Principal Software Engineer to lead the AI Security Platform for code analysis across enterprise codebases. The role emphasizes platform ownership, architectural evolution, and scalable security automation in a cloud-native environment....
$130k - $190k
...consensus. The Team: As AI tools drive immense value across... ...cost curve. The unified AI Platform Team builds the centralized software... ...are looking for a Platform Engineer to help build and... ...alerts. Partner with the Inference Infrastructure pod to ensure...Full timeTemporary workFlexible hours- ...The Chamberlain Group is seeking a Sr. Manager of Software Engineering to lead the Builder Tools and Governed Agent Platform, guiding Forward Deployed Engineers and ensuring AI-driven automation aligns with guardrails. Hybrid work in Oak Brook, IL offers leadership...
$120k - $140k
...technology across the enterprise, and our cloud platforms are central to how our customer-facing... .... As a Senior DevOps & Site Reliability Engineer, you will be hands-on at the center of... ...Dev, QA, UAT, and Production environments.AI-Enabled DevOps & Platform AutomationUse...Permanent employmentTemporary workWork experience placementH1bLocal areaRemote workFlexible hours$106.08k - $126.88k
...demands to travel. Key Responsibilities: As a Release Train Engineer, you will be responsible for facilitating Agile Release Train... ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and...Hourly payFull timeLive inWork at officeLocal areaImmediate startFlexible hoursShift work- ...to own the Builder Tools vision and guide the governed Agent Platform, including Cloud Agents, per-agent identity, and a Cortex index. You will lead Forward Deployed Engineers to ensure alignment with our AI-driven Builder Tools guardrails. Responsibilities include...
- ...advisors, and investment professionals.As a Quantitative Software Engineer - Research Platform, you will help shape Schwab’s research technology ecosystem... ....Experience applying machine learning, generative AI, or advanced analytics capabilities to business problems.Ability...Full timeWork at office
$63 - $90 per hour
...services provider, is seeking a Senior DevOps Engineer / Site Reliability Engineer (SRE) to... ...support enterprise backup, cyber recovery, and platform resiliency initiatives. The ideal... ...for this job, you agree to receive calls, AI-generated calls, text messages, or emails...Permanent employmentContract workLocal area$164.6k - $288k
## Sr Principal Software Engineer, AI Security PlatformApply: Chicago, IL: Full time: Posted Today: R161850**About Northern Trust** As a... ...technical development of our AI-powered Code Security Scanner platform. This platform leverages frontier AI models and advanced security...Full timeH1bWorldwideFlexible hours$209k - $238.5k
...Sr. Manager, Software Engineering, Back End (Enterprise Platforms Technlogy) Do you love building and pioneering in the technology space? Do you enjoy solving... ...and decisioning 3+ years of experience in implementing AI and Gen AI for marketing use cases 7+ years of people...Full timePart timeInternshipLocal area$116.2k - $229.1k
Position Summary AI & Engineering/EaaS - DevOps Engineer IIIBuild and scale modern DevOps capabilities that help clients accelerate... ..., implementation, and optimization of cloud, automation, and platform engineering solutions across complex environments. This role...Local area- ...IllinoisHybridFull Time$160k - $190kSenior DevOps Engineer An established industrial supply company... ...Engineer to help expand and enhance its Platform Product team. This hybrid role is located... ..., and observability Utilize a pro-AI workplace mindset with the ability to quickly...Full time
- ...customer outcomes within strategic accounts from deployment to adoption. You will design, build, and ship capabilities on Axon's platform onsite, working across agents, workflows, and data tools to deliver production-ready solutions. You'll lead engagements with agency...
$120k - $140k
...technology across the enterprise, and our cloud platforms are central to how our customer-facing... .... As a Senior DevOps & Site Reliability Engineer , you will be hands-on at the center of... ..., UAT, and Production environments. AI-Enabled DevOps & Platform Automation...Permanent employmentFull timeTemporary workWork experience placementH1bLocal areaRemote workFlexible hours- ...Job Title: Databricks Data Engineer ( Databricks, AWS/Azure, Python, PySpark, UDP) - Chicago Location: Chicago, IL Secondary... ...development (production-level code) ~ UDP or modern data platform architecture experience ~ Ability to work hybrid model in Chicago...Relocation
- ...Job Description Job Description Senior Veritas eDiscovery Platform (eDP) Engineer Employment Type: Full-Time, Executive-Level Department... ...cgsfederal.com #CJ We may use artificial intelligence (AI) tools to support parts of the hiring process, such as...Full timeFor contractorsRemote workFlexible hours
- ...skilled and experienced Senior/Staff QA Engineer to support a critical customer engagement... ...automated coverage for a fast-moving agentic AI product.The successful candidate will... ...to-end (E2E) testing against a live agent platform.Autonomy: Ability to independently define...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Inference Platform Engineer. Be the first to apply!
- senior ai engineer Chicago, IL
- machine learning ai engineer Chicago, IL
- ai ml engineer Chicago, IL
- ai developer Chicago, IL
- ai prompt engineer Chicago, IL
- ai engineer Chicago, IL
- ai engineer remote Chicago, IL
- client platform engineer Chicago, IL
- platform engineering manager Chicago, IL
- senior platform engineer Chicago, IL





