Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Platform Engineer

$200k - $250k
Full-time

DRW

DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk.

Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.

We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters–it's how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.

About the Role

We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use.

You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy. You'll own the serving platform end-to-end: onboarding newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet.

What You'll Do

  • Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes.
  • Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems.
  • Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies.
  • Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models.
  • Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing.
  • Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads.
  • Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
  • Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments.
  • Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements.
  • Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads.

What We're Looking For

The Tech

  • Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions.
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.
  • Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching.
  • Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes.
  • Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode.
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads.
  • Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers.
  • Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness.

The Intangibles

  • You take a measurement-driven approach to performance optimization.
  • You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries.
  • You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems.
  • You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes.
  • You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance.
  • You communicate clearly and can explain complex performance tradeoffs across engineering teams.

The annual base salary range for this position is $200,000 to $250,000 depending on the candidate’s experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus. In addition, DRW offers a comprehensive suite of employee benefits including group medical, pharmacy, dental and vision insurance, 401k (with discretionary employer match), short and long-term disability, life and AD&D insurance, health savings accounts, and flexible spending accounts.

For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at

California residents, please review the California Privacy Notice for information about certain legal rights at

[#LI-VD1]

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Inference Platform Engineer in Chicago, IL vacancy
  • $165k - $225k

     ...Moonlite delivers high-performance AI infrastructure for...  ...software-defined networking (SDN) platform that enables high-performance...  ...distributed computing, model training, inference, and data-intensive workloads...  ...– enabling researchers and engineers to access enterprise-grade... 
    Suggested
    Immediate start
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    28 days ago
  • $165k - $225k

     ...Moonlite delivers high-performance AI infrastructure for organizations...  ...out our GPU-accelerated compute platform that powers distributed AI training and inference, large-scale simulations, and computational...  ...-enabling researchers and engineers to programmatically access high-... 
    Suggested
    Immediate start
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    28 days ago
  •  ...Travel: Occasional international travel required Job Duration: Long Term Contract Job Summary We are seeking an AI Platform Engineer to build and scale enterprise AI platform capabilities that enable multiple product teams. This role focuses on developing production... 
    Suggested
    Long term contract
    Remote work

    Stellar IT Solutions

    Chicago, IL
    19 hours ago
  • $165k - $225k

     ...Moonlite delivers high-performance AI infrastructure for...  ...comprehensive infrastructure platform that bridges our physical infrastructure...  ...for large-scale computation, inference, simulations, and training....  ...that enable researchers and engineering teams to programmatically... 
    Suggested
    Immediate start
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    12 days ago
  • $119.4k - $204.6k

    Enterprise Platform Strategy Lead Essential Responsibilities: Define enterprise-wide platform strategy, vision...  ...standards and guardrails across all platform engineering domains. Drive innovation in cloud, data, and AI platforms (Fabric, OpenAI/Foundry, etc.). Lead large... 
    Suggested
    Full time
    Temporary work
    Part time
    Work from home
    3 days per week

    Alliant

    Chicago, IL
    2 days ago
  • $133.9k - $154.5k

     ...solutions and opportunities for our customers and stakeholders. Every day we work toward transforming global markets. The AI Platform Engineering Lead drives the AI Platform Operations team, guiding platform strategy, governance, and stakeholder engagement. They align... 
    Full time
    Contract work
    Temporary work
    Flexible hours

    Intercontinental Exchange

    Chicago, IL
    5 days ago
  • $124.36k - $146.3k

     ...from Day One. Job Description Job Summary The Senior Engineer (Generative AI) is responsible for designing, developing, and deploying...  ...observability practices for production environments 3. Cloud, Platform & Scalability Engineering Develop and deploy GenAI... 
    Full time
    Temporary work
    Work experience placement
    Local area
    3 days per week

    U.S. Bank

    Chicago, IL
    6 hours ago
  • $119k - $169.4k

     ...While the internal title for this position is Senior Automation Engineer , this role has been posted externally under a different title to...  ...intelligence. This role sits at the intersection of advanced AI engineering, security architecture, and organizational strategy... 
    Work at office
    Immediate start

    Cboe Global Markets

    Chicago, IL
    3 days ago
  • $130k - $190k

     ...consensus. The Team: As AI tools drive immense value across...  ...cost curve. The unified AI Platform Team builds the centralized software...  ...are looking for a Platform Engineer to help build and...  ...alerts.  Partner with the Inference Infrastructure pod to ensure... 
    Full time
    Temporary work
    Flexible hours

    Drw

    Chicago, IL
    21 days ago
  • $148.2k - $200.85k

     ...of the financial services industry's most demanding SaaS platforms. Our Platform Engineering team sits at the intersection of DevOps, private and public...  ...environments. Demonstrated comfort and fluency with AI-assisted workflows - actively leverages AI tools (e.g., Claude... 
    Permanent employment

    Clearwater Analytics

    Chicago, IL
    3 days ago
  • $145k

     ...Sydney, Shanghai, Singapore, and London. What you'll do as a Platform Engineer Intern at Akuna: We are seeking Platform Engineer Interns...  ..., you may apply to any Trader roles of interest. We are an AI-friendly company and encourage employees to leverage AI in their... 
    Summer work
    Internship
    Work at office

    Akuna Capital

    Chicago, IL
    1 day ago
  • $147.76k - $240.11k

     ...assets worldwide, our teams use data, technology, advanced analytics, and AI capabilities to help our customers build a better, more sustainable world. Job Summary The PIM Platform Engineering Manager leads engineering strategy and delivery for the PIM platform,... 
    Full time
    Part time
    Worldwide
    Relocation package
    Flexible hours

    Caterpillar Inc.

    Chicago, IL
    4 days ago
  • $162k - $262k

    84.51° is looking for a Director of Software Engineering to lead its AI Enablement team. This role involves designing and maintaining the enterprise's agentic backbone, providing technical leadership, and managing both onshore and offshore teams. The ideal candidate will... 

    84.51°

    Chicago, IL
    6 days ago
  • $112.8k - $153.7k

    Senior Engineering Role At Armanino, you determine your career path. This means it's possible...  ...improve cloud data and analytics platforms (e.g., Microsoft Fabric, Snowflake, Databricks...  ...Databricks, etc.) are a plus. Hands-on AI/ML/GenAI enablement experience (model lifecycle... 
    Contract work
    Local area
    Flexible hours

    Armanino

    Chicago, IL
    6 days ago
  • $63 - $90 per hour

     ...services provider, is seeking a Senior DevOps Engineer / Site Reliability Engineer (SRE) to...  ...support enterprise backup, cyber recovery, and platform resiliency initiatives. The ideal...  ...for this job, you agree to receive calls, AI-generated calls, text messages, or emails... 
    Permanent employment
    Contract work
    Local area

    KellyMitchell Group

    Chicago, IL
    2 days ago
  • Senior Software Engineer Welcome to Our World We've been leading the charge in the affiliate...  ...the largest, most reliable partnership platforms with impeccable, personalized service. Founded...  .... We are committed to finding out how AI can amplify our productivity. We view AI... 

    Digitas

    Chicago, IL
    4 days ago
  •  ...Job Description Job Description Senior Veritas eDiscovery Platform (eDP) Engineer Employment Type: Full-Time, Executive-Level Department...  ...cgsfederal.com #CJ We may use artificial intelligence (AI) tools to support parts of the hiring process, such as... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Contact Government Services, LLC

    Chicago, IL
    29 days ago
  • $120k - $140k

     ...technology across the enterprise, and our cloud platforms are central to how our customer-facing...  .... As a Senior DevOps & Site Reliability Engineer , you will be hands-on at the center of...  ..., UAT, and Production environments. AI-Enabled DevOps & Platform Automation... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    H1b
    Local area
    Remote work
    Flexible hours

    Berlin Packaging

    Chicago, IL
    3 days ago
  • $209k - $238.5k

     ...Sr. Manager, Software Engineering, Back End (Enterprise Platforms Technlogy) Do you love building and pioneering in the technology space? Do you enjoy solving...  ...and decisioning 3+ years of experience in implementing AI and Gen AI for marketing use cases 7+ years of people... 
    Full time
    Part time
    Internship
    Local area

    Capital One National Association

    Chicago, IL
    4 days ago
  •  ...Job Title: Databricks Data Engineer ( Databricks, AWS/Azure, Python, PySpark, UDP) - Chicago Location: Chicago, IL Secondary...  ...development (production-level code) ~ UDP or modern data platform architecture experience ~ Ability to work hybrid model in Chicago... 
    Relocation

    Innosystech Inc

    Chicago, IL
    2 days ago
  • $189k - $236k

     ...Learn more about our Total Rewards philosophy . AI is a fundamental part of how work gets done at Gusto....  ...process. About the Role: We’re hiring seasoned engineers to join our teams that work on core platform capabilities, improving our existing systems for... 
    Remote job
    Full time
    Work at office
    Local area
    2 days per week
    3 days per week

    gusto

    Chicago, IL
    more than 2 months ago
  • $200k - $250k

     ...We're looking for a senior level DevOps Engineer to join a small, high-impact team that builds...  ...at scale, a firm-wide observability platform, CI/CD infrastructure, and workflow orchestration...  ..., storage systems. Thoughtful use of AI coding assistants and interest in AI-... 
    Worldwide
    Flexible hours

    DV Trading

    Chicago, IL
    28 days ago
  • DevOps Engineer All IT Solutions Chicago, IL Hybrid work Contract Shift and schedule We are currently seeking a qualified DevOps Engineer...  ...closely matches the position will be contacted directly by the AIS recruiting team. Job Summary All IT Solutions is seeking a... 
    Contract work
    Local area
    Shift work
    Weekend work

    All IT Solutions

    Chicago, IL
    13 days ago
  • $40 - $50 per hour

     ...investment adviser. Overview: As a DevOps Engineer Intern, you'll work directly within a...  ...that owns firm-wide infrastructure and platform solutions. This is a high-impact, high-...  ...cutover planning). Contribute to internal AI-assisted tooling: use AI thoughtfully (... 
    Hourly pay
    Summer work
    Internship
    Worldwide

    DV Trading

    Chicago, IL
    20 days ago
  • $140k - $180k

     ...We are seeking an experienced DevOps Engineer to lead and enhance our organization’s technical...  ...Strong understanding of IT operations, platform services, and system reliability...  ...growing We may use artificial intelligence (AI) tools to support parts of the hiring process... 
    Flexible hours

    Supernova Technology™

    Chicago, IL
    2 days ago
  •  ...security requirements. Support AI and machine learning...  ...experience with Google Cloud Platform services. Strong problem-solving...  ...Professional Cloud Architect, Data Engineer). Familiarity with Palo...  ...interviewing at The Judge Group by 2x Inferred from the description for this... 
    Contract work
    Internship
    3 days per week

    The Judge Group

    Chicago, IL
    2 days ago
  • EY is seeking a data and AI platform leader to drive delivery across client engagements in a hybrid model in the United States. You will guide cross-functional teams of engineers, data scientists, and designers to build scalable data platforms and intelligent solutions... 

    EY

    Chicago, IL
    3 days ago
  • $164.5k - $197.5k

    Job ID 108450 Work Areas Technology & Engineering Employment Type Permanent Full-Time Location...  ...Solutions Group (TSG) , partnering closely with Platform Infrastructure, Architecture Center of...  ...portfolio includes cloud infrastructure, AI tooling, and self-service developer... 
    Permanent employment
    Full time
    Work at office
    1 day per week

    Bain & Company

    Chicago, IL
    6 days ago
  • $122.4k - $228k

     ...passionate professional for a Senior Cloud, AI & Data Security Engineer role who wants to design and...  ...services across AWS, Azure, and AI/ML platforms. We need someone who can establish the...  ...feature stores, training environments, and inference endpoints Knowledge of Generative... 
    Part time
    Local area

    Koitecc Solutions

    Chicago, IL
    19 hours ago
  • $314.8k - $359.3k

    Senior Staff Full-stack Engineer - Customer Platforms As a Senior Staff Engineer at Capital One, you will be part of a community of technical experts...  ...~ Expertise in the strategic application of agentic AI coding tools to solve systemic development challenges... 
    Full time
    Part time
    Local area

    Capital One

    Chicago, IL
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Platform Engineer. Be the first to apply!