Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer

F5 Networks Inc

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices—with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy. This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance. Key ResponsibilitiesHigh-Performance AI ServingBuild and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale. Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications. Hardware Acceleration and OptimizationProfile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs. Collaborate with hardware teams to maximize utilization and performance across various computational environments. Inference Orchestration and ScalabilityDesign and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability. Performance Monitoring and ObservabilityEstablish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs). Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale. Technical RequirementsRequired Skills:Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows. Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization. Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure. Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs. Preferred Experience:Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention. Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels). Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges. Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts. Success Metrics (KPIs):Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second. Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization. Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments. Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput. Why Join F5?F5 empowers you to push boundaries in AI optimization and high-performance engineering. Joining our team means: Collaborating with cutting-edge technologies and hardware solutions to support real-time AI applications. Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving. Driving transformative projects that deliver real-time AI reliability to global customers while maintaining cost and efficiency standards. Working on advanced MLOps solutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment. What Success Looks Like:As an AI Inference Engineer at F5, success is measured by your ability to: Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems. Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement. Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals. Translate complex AI and inference workflows into practical solutions that align with F5's strategic objectives. The base pay range per annum for this position is: $176,600 - $265,000F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change. You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: F5 reserves the right to change or terminate any benefit plan without notice.#LI-ZB1The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com View email address on click.appcast.io).Equal Employment OpportunityIt is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting View email address on click.appcast.io: San Jose; SeattleType: Full time

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer in Seattle, WA vacancy
  • $151.8k - $332.2k

    What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on... 
    Suggested
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    12 hours ago
  • $188k - $275k

     ...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,...  ...'ll Do Description of the team: The Inference team is responsible for delivering high-performance...  ...: We are looking for an Applied AI Engineer to help us understand, measure, and... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Seattle, WA
    5 days ago
  • NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language... 
    Suggested

    NVIDIA

    Seattle, WA
    3 days ago
  •  ...As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient...  ...CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution...  ...latency and maximize memory bandwidth on AI accelerators. Write production-level,... 
    Suggested
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    21 days ago
  • $154.56k - $193.2k

     ...hyperscaler for the edge, delivering modular AI infrastructure from first deployment to...  ...Armada is seeking exceptional AI Engineers to build and deploy intelligent systems at...  ...analysis, anomaly detection, or distributed AI inference. This role is intended for engineers... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    Armada

    Bellevue, WA
    4 days ago
  •  ...Senior Staff AI Engineer The AI Scientist will work in teams addressing statistical, machine learning and data understanding problems...  ...clinical reports, and other healthcare datasets. Develop robust inference pipelines, model serving infrastructure, and AI services... 
    Worldwide
    Flexible hours

    GE

    Bellevue, WA
    3 days ago
  • $193.3k - $261.5k

     ...cloud-scale machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance...  ...in production on GPUs, AWS Neuron, TPUs, or other AI accelerator hardware- Experience extending or... 
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    12 hours ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...tools Experience investigating, and resolving, training & inference performance end to endDebugging and optimization... 
    Full time
    Remote work

    Nvidia

    Seattle, WA
    12 hours ago
  • $152k - $241.5k

     ...recently, GPU deep learning ignited modern AI — the next era of computing — with the...  ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep...  ...-based AI compiler that powers NVIDIA’s inference engine end to end, with a focus on... 
    Full time
    Remote work

    Nvidia

    Seattle, WA
    2 days ago
  • $151.8k - $332.2k

     ...What You Can ExpectYou'll design, implement, and own the inference systems that serve Zoom's AI models at production scale -- across real-time...  ...implementationsDrive technical design and set the bar for inference engineering practices across the teamWhat We're Looking ForA... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    12 hours ago
  • $163.2k - $220.8k

     ...and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI Security...  ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM access... 
    Full time
    Work experience placement
    Remote work
    Worldwide
    Shift work

    Wilson Sonsini Goodrich & Rosati

    Seattle, WA
    3 days ago
  • Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own...  ...(OpenAI-style APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)Vector DB Ideal Experience... 

    Spectraforce Technologies

    Seattle, WA
    1 day ago
  • $151.8k - $332.2k

    What you can expectWe are seeking an experienced AI Infrastructure Engineer to join our AI Incubation team. You will be focused on building and...  ...workflowsCollaborating with AI researchers to implement efficient training and inference pipelinesWhat we’re looking forHave a bachelor's degree in... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    1 day ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions...  ...Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Seattle, WA
    2 days ago
  •  ...Senior AI Engineer – Privacy Location: Bellevue WA Must have skills – skill 1 – 7yrs of exp – AI Engineer – Privacy skill 2 – 7yrs...  ...Snowflake, or PySpark to support AI model training, fine-tuning, and inference. Apply prompt engineering, few-shot learning, and fine-... 

    Software Technology Inc

    Bellevue, WA
    1 day ago
  • $176.76k - $232k

     ...people.About this team The Enterprise Data & AI team is a strategic and operational...  ....Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable AI...  ...architectures and system design for serving AI/ML inference solutions in production. You will help... 
    Permanent employment
    Full time
    Contract work
    Part time
    Work visa

    Lululemon Athletica

    Seattle, WA
    3 days ago
  • $220k - $292k

     ...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...TPUs, or custom ASICs) for training and inference workloads. Model Observability at Scale:... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  •  ...the forefront of a new era in enterprise AI — one defined not by model capability...  ...of frontier AI research and production engineering — investigating the foundational challenges...  ...persistence architectures, model selection and inference routing strategies, autonomy and goal-... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Relocation

    Accenture

    Seattle, WA
    2 days ago
  • $143.7k - $194.4k

     ...commerce.Advertiser Growth Tech (AGT) is an engineering team with the mission to enhance the...  ...scalable code. You will work on Agentic AI initiatives, contributing to systems that...  ..., model training, optimization, and inference infrastructure. Lead technical design discussions... 
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  •  ...Lead AI Engineer in the Platforms and Products ZS is a place where passion changes lives. As a management consulting and technology...  ...retrieval-augmented generation, fine-tuning workflows, and scalable inference pipelines. Design and implement LLM-powered applications... 
    Local area
    Work from home
    Worldwide
    Flexible hours

    ZS Associates

    Bellevue, WA
    4 days ago
  • $165k - $242k

     ...degradation, rollback/traffic-shift strategies. Mentor IC1/IC2 engineers; review cross-team designs and elevate coding/testing...  ...(Prometheus, Grafana, OpenTelemetry). Practical knowledge of inference internals: batching, caching, mixed precision (BF16/FP8), streaming... 
    Permanent employment
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    12 hours ago
  • $139k - $204k

    What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to...  ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    12 hours ago
  • TikTok is seeking a Research Engineer/Scientist to design and implement efficient large-scale generative AI models, focusing on distillation and compression to enable scalable...  ...model acceleration, hardware-efficient inference, and transferring capabilities from foundation... 

    TikTok

    Seattle, WA
    12 hours ago
  • $91.1k - $179.5k

     ...significant impact on our clients’ success. We are hiring an AI Engineer to build and operate the data, features, and GenAI foundations...  ...pipelines and services that support model training, real-time inference, and LLM applications using Claude-, GPT/Codex-, and Gemini-class... 
    Local area

    Deloitte

    Seattle, WA
    4 days ago
  • $171k - $240k

     ...gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual...  ...resources, and support you need to grow your career.AI at BrexAI Engineering at Brex is redefining how businesses run their finances by building... 
    Work at office
    Remote work
    Work from home
    Shift work

    Brex

    Seattle, WA
    12 hours ago
  • $177.1k - $387.5k

     ..., and own the platform that powers Zoom AI Services, enabling AI capabilities to be...  ..., cloud infrastructure, and AI platform engineering to build reliable, high-performance services...  ..., including model serving, AI inference platforms, GPU/CPU resource management,... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    12 hours ago
  •  ...promised to change everything across a business—until generative AI. Today, AI is the number one driver of business reinvention. And...  ...and help lead that change.You Are:As a Snowflake Advanced AI Engineer, you will design, build, and operationalize artificial intelligence... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Shift work

    Accenture

    Seattle, WA
    3 days ago
  • $342.7k

     ...strategy for Splunk and Cisco, and we are building best-in-class AI capabilities into our platform, security and observability...  ...journey to digital resilience. We are looking for a Distinguished Engineer with outstanding technical leadership in the AI/ML area to drive... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Seattle, WA
    3 days ago
  •  ...with solutions created by the #1 company in e-signature and contract lifecycle management (CLM). What you'll do As an AI Agentic Engineer on the Platform AI Engineering team, you will drive the next wave of automation in IT operations by designing and deploying agentic... 
    Permanent employment
    Full time
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    Seattle, WA
    12 hours ago
  • $130k - $170k

    CompanyFounded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the digital economy. Taxbit's AI-enabled platform streamlines compliance related to digital assets, payments, and other financial transactions. Its... 
    Work at office
    Work from home

    Taxbit

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer. Be the first to apply!