SRE, AI Inference Engineer
F5 Networks Inc
At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation. Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices—with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy. This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance. Key ResponsibilitiesHigh-Performance AI ServingBuild and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale. Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications. Hardware Acceleration and OptimizationProfile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs. Collaborate with hardware teams to maximize utilization and performance across various computational environments. Inference Orchestration and ScalabilityDesign and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration. Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability. Performance Monitoring and ObservabilityEstablish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs). Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale. Technical RequirementsRequired Skills:Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows. Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization. Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure. Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs. Preferred Experience:Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention. Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels). Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges. Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts. Success Metrics (KPIs):Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second. Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization. Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments. Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput. Why Join F5?F5 empowers you to push boundaries in AI optimization and high-performance engineering. Joining our team means: Collaborating with cutting-edge technologies and hardware solutions to support real-time AI applications. Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving. Driving transformative projects that deliver real-time AI reliability to global customers while maintaining cost and efficiency standards. Working on advanced MLOps solutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment. What Success Looks Like:As an AI Inference Engineer at F5, success is measured by your ability to: Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems. Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement. Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals. Translate complex AI and inference workflows into practical solutions that align with F5's strategic objectives. The base pay range per annum for this position is: $176,600 - $265,000F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5’s differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change. You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5’s benefits can be found at the following link: F5 reserves the right to change or terminate any benefit plan without notice.#LI-ZB1The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com View email address on us.fitly.work).Equal Employment OpportunityIt is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting View email address on us.fitly.work: San Jose; SeattleType: Full time
$151.8k
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on...SuggestedWork at officeRemote work- ...industry. Our mission is to enable every engineering organization to build its own self-... ...design workforce. Agentrys Studio combines AI agents, engineering knowledge, agent-... ...AI Engineer to own the model training, inference, and infrastructure that power Agentrys'...SuggestedFull time
- Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product...Suggested
$106.9k - $176.5k
...world. The opportunity We are seeking AI Systems Engineers to own the security and trust fabric of... ...Enterprise Security / Cloud Platform / SRE to ensure independent review, alignment... ...specific trust challenges of confidential AI inference (models/secrets inside enclaves)....SuggestedWork experience placementSummer holidayRemote workFlexible hours$160.08k - $240.12k
...Job Description Summary The AI Scientist will work in teams addressing statistical... ...HealthCare is seeking a Senior Staff AI Engineer to design, build, deploy, and optimize production... ...healthcare datasets. Develop robust inference pipelines, model serving infrastructure,...SuggestedWorldwideVisa sponsorshipWork visaRelocation packageFlexible hours$190k - $225k
...AssemblyAI builds the best-in-class Voice AI models powering the next generation of voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio... ...About the role: We're hiring a Software Engineer to help turn cutting-edge AI research...$92k - $135k
...Description CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ...more at What You'll Do: Join the Inference team to ship production features that improve... ...quickly with mentorship from experienced engineers. About the role: Implement well-...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... ...hands-on experience delivering agentic AI to production, beyond LLM wrappers or... ...domains. Preferred knowledge includes SRE concepts such as SLOs/SLIs, error budgets...Full time$165.2k - $223.6k
...you passionate about building infrastructure that powers AI at scale? Do you want to work on systems that serve millions... ...(ALICE) team is looking for Software Development Engineers to join our Inferences Services Team.We build products and solutions that enable...InternshipLocal areaFlexible hours$152k - $241.5k
...recently, GPU deep learning ignited modern AI — the next era of computing — with the... ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep... ...-based AI compiler that powers NVIDIA’s inference engine end to end, with a focus on...Full timeRemote work- ...positive impact in critical industries through AI transformation. We specialize in physics-... ...anchors on two systems. The first is our inference control plane — open-weight models and... ...before you go home — building the fork engine, the guest agent, and the multi-substrate...Full timeRemote workWork visaFlexible hoursDay shift
$176.76k - $232k
...people.About this team The Enterprise Data & AI team is a strategic and operational... ....Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable AI... ...architectures and system design for serving AI/ML inference solutions in production. You will help...Permanent employmentFull timeContract workPart timeWork visa$143.7k - $194.4k
We are looking for an **AI/ML Engineer** to build, deploy, and operate the ML/AI systems that power the agentic decision intelligence workflow... ..., rollback, and canary deployment. Ensure SLA compliance for inference latency and availability- Own the operational health of AI/ML...InternshipFlexible hours$130k - $170k
...Company Founded by CPAs, tax attorneys, and engineers, Taxbit is the leading innovator automating global tax reporting for the digital economy. Taxbit's AI-enabled platform streamlines compliance related to digital assets, payments, and other financial transactions....Full timeWork at officeWork from homeFlexible hours$188k - $275k
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers... ...2025. Learn more at What You'll Do: Inference Platform Team The Inference team builds... .... About the role: As a Staff Software Engineer (IC5) on the Inference team, you will act...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$209.1k - $282.9k
Software Engineer AI Inference Runtime Team As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models. You will lead hands-on work across scheduling, batching...Work at officeLocal area- Job Title: Applied AI EngineerLocation: 100% RemoteFull-Time Role Role As an Applied AI Engineer, you will turn model capabilities into real product behavior.You will own... ...(OpenAI-style APIs, LLaMA, Qwen, etc.)Inference / serving (e.g. vLLM)Vector DB Ideal Experience...
- Senior Lead - AI Engineering ZS is a place where passion changes lives. As a management consulting and technology firm focused on improving... ...Generation (RAG), fine-tuning workflows, and scalable inference pipelines. Design and implement LLM-powered applications using...Local areaWork from homeWorldwideFlexible hours
- AI Training Infrastructure Engineer Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) About the Opportunity A... ...including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by...Work at officeRelocation3 days per week
$220k - $292k
...of systems is powered by Lattice OS, an AI-powered operating system that turns thousands... ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems... ...TPUs, or custom ASICs) for training and inference workloads.Experience building production...Full timeWork experience placementImmediate start$145k
...Agentic AI Engineer Envorso Sports · Seattle, WA (Hybrid) or Remote, US · Full-time About Envorso Sports Envorso Sports builds the platform that runs a professional sports team's entire fan-facing business, ticketing, website, and mobile app, from one system...Full timeLive inRemote workVisa sponsorship- ...Job Title: AI Engineer Location- Bellevue (Seattle area), WA ( Day 1 Onsite Job Type: Contract # What is the overall exp and no. of years exp in GenAI, Agentic AI work? # Expectation is overall experience of 10+ and around 3 yrs in GenAI / Agentic AI...Contract workTemporary work
$118.9k - $202.1k
...Learn how Premera supports our members, customers and the communities that we serve through our Healthsource blog: . As an AI Engineer III , you will work within a team responsible for taking AI/ML solutions from initial concept to full production implementation....Full timeRemote workWork from home- ...About the Role At Ravenna, we are looking for experienced engineers with a passion for AI and a strong track record of building and shipping production systems. Our team is combining the latest advances in large language models with proven machine learning techniques...Flexible hours
- ...technology products. As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate... ..., scalable cloud platforms optimized for AI/ML workloads. Partner with AI teams to... ...architecture, ML training, and inference. Experience with Infrastructure as Code...For contractors
$320k
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole...Work at officeWorldwideVisa sponsorshipFlexible hours$61k - $101k
...formal training or certification in software engineering concepts, along with 5+ years of applied... ...transformer architecture, training, and inference. We require experience with... ...the effective use of enterprise-approved AI-assisted software development tools, with...Full timeFor contractors- ...About the Role Socure is looking for an AI Enablement Engineer to accelerate AI adoption across engineering by turning standards, tooling patterns, and proven workflows into repeatable day-to-day practice. This role sits at the intersection of AI tooling enablement...Work at officeImmediate start
$144.7k - $261.3k
...environments, cloud infrastructure, and ML/AI GPU platforms for AV research and... ...Role GM is looking for a Senior Performance Engineer to join the AV Capacity and Performance Engineering... ...of large-scale ML training and inference environments. Your skills & abilities (Required...Work at officeLocal areaRemote workWork from homeFlexible hours3 days per week- ...AI Engineer Location: Plano, TX 75024 / Indianapolis, IN 46240-3797/ Portsmouth, NH 03801/ Boston, MA 02116 / Seattle, WA - 98154 FULLTIME ONLY Job Description Must Have Technical/Functional Skills • Required Technical Skill Set: Python, REST APIs, GitHub...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE, AI Inference Engineer. Be the first to apply!
- site reliability engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- senior ai engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- ai ml engineer Seattle, WA
- ai developer Seattle, WA
- ai prompt engineer Seattle, WA
- ai engineer Seattle, WA
- ai engineer remote Seattle, WA
- junior site reliability engineer



