AI Infrastructure Engineer
Sciforium
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications. About the Role We are looking for an AI Infrastructure Engineer to own the entire software stack of our GPU clusters — from kernel tuning and GPU drivers up through schedulers, containers, and ML frameworks. While our Hardware Operations team keeps the physical machines healthy and connected, you define what a production-ready node looks like in software: you author the images, playbooks, and pipelines that take a freshly provisioned server to a fully validated GPU node, and you keep the fleet consistent, upgradable, and fast. You will serve two demanding customer groups — our foundation model training teams and our model serving/product teams — ensuring both run on correctly configured, well-managed, high-performance infrastructure. Key Responsibilities OS Bring-Up & Node Lifecycle Engineering Golden Images & Automated Bring-Up: Own the node software definition — versioned OS images, kernel tuning (NUMA, hugepages, IRQ affinity, cgroups), GPU/NIC driver stacks — and the automated pipeline that takes a node from base OS to production-ready. Validation & Burn-In: Build automated acceptance suites (DCGM diagnostics, nccl-tests/RCCL tests, bandwidth and topology checks, HPL) that gate every node before it enters a scheduler pool. Fleet Maintenance: Execute rolling kernel/driver/toolkit upgrades with minimal disruption to running workloads; enforce configuration consistency, detect drift, and maintain the driver ↔ CUDA/ROCm ↔ framework compatibility matrix across the fleet. Self-Healing Operations: Automate detection of unhealthy nodes (Xid/ECC errors, link flaps, thermal throttling), with cordon/drain/reboot/re-image workflows and clean handoff to Hardware Operations for physical repair or RMA. Configuration Management & Automation Infrastructure as Code: Manage all node and cluster configuration through Ansible/SaltStack playbooks in Git, with peer-reviewed changes, CI validation, and canary rollouts before fleet-wide deployment. Provisioning Pipelines: Build and maintain image/provisioning tooling (PXE, MaaS, Packer, or similar) so new or re-imaged nodes are reproducible, not hand-crafted. Operational Tooling: Develop Python/Bash tooling for cluster operations, health reporting, and workflow automation. Orchestration & Scheduling (Kubernetes & Slurm) Kubernetes for Serving: Deploy and operate GPU-enabled Kubernetes for inference workloads — NVIDIA GPU Operator, device plugins, node feature discovery, topology-aware scheduling, and MIG/MPS partitioning where appropriate. Training Schedulers: Operate Slurm (or Run:AI) for multi-node training — partitions, QoS, preemption, accounting, and container integration (enroot/pyxis). Container Platform: Maintain base images, registries, and the NVIDIA Container Toolkit / ROCm container stack; keep training and serving images lean, current, and reproducible. GPU Driver & ML Stack Engineering Driver & Runtime Lifecycle: Build, deploy, and debug the full accelerator stack — NVIDIA (CUDA toolkit, cuDNN, NCCL, Fabric Manager) and AMD (ROCm, RCCL) — including kernel modules (DKMS), GPUDirect RDMA/Storage, and the RDMA software stack (MOFED/DOCA). Framework Environments: Maintain curated, optimized PyTorch and JAX environments with sane dependency and version management for researchers and production services. Distributed Performance: Tune NCCL/RCCL across NVLink/NVSwitch and InfiniBand/RoCE fabrics, ensure topology-aware job placement, and run continuous communication/throughput benchmarks to catch regressions. Advanced Debugging & Observability Escalation Point: Own the hard problems — NCCL hangs and timeouts, CUDA memory leaks, ROCm kernel crashes, straggler nodes, and unexplained throughput drops. Observability: Own software-layer monitoring (DCGM exporter, Prometheus/Grafana, alerting) plus job-level GPU utilization and cluster efficiency reporting. Qualifications Must-Haves: 5+ years in systems/infrastructure engineering with significant GPU cluster, HPC, or large-scale ML infrastructure experience. Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field. Deep Linux internals expertise: kernel modules/DKMS, systemd, cgroups, NUMA, and system performance tuning. Hands-on experience with NVIDIA (CUDA) and/or AMD (ROCm) driver and runtime stacks on modern accelerators (H200/B200, MI325x/MI355x class), including kernel-level debugging. Production Kubernetes experience with GPU workloads, plus working knowledge of HPC schedulers (Slurm/Run:AI) — or the reverse (deep Slurm, working K8s). Strong configuration management experience (Ansible or SaltStack) with Git-based, code-reviewed infrastructure workflows. Provisioning and image tooling experience (Packer, MaaS, Foreman, Terraform, or similar) for automated, reproducible node builds. Client-side experience with distributed filesystems (Lustre, GPFS, Weka) and checkpoint I/O optimization. Container fluency: Docker/containerd and the NVIDIA Container Toolkit or ROCm equivalent. Proficiency in Python and Bash for automation and tooling. Working knowledge of NCCL and RDMA networking (InfiniBand/RoCE, GPUDirect) and of PyTorch/JAX runtime behavior. Nice-to-Haves: Experience directly supporting foundation model training teams — multi-node job failure debugging, checkpoint pipeline tuning, and framework-level performance triage — ideally in a startup or research-heavy environment. Experience deploying and tuning inference/serving stacks (vLLM, Triton Inference Server, TensorRT-LLM) for latency and throughput targets. GPU/system profiling tools: Nsight Systems/Compute, rocprof, perf, eBPF. Benefits include Medical, dental, and vision insurance 401k plan Daily lunch, snacks, and beverages Flexible time off Competitive salary and equity Equal opportunity Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
$190k - $310k
The role As an AI platform engineer, you'll build the products, interfaces, and tools that define how people interact with Applied Compute... ...for the enterprise. We provide the continual learning infrastructure for companies to build agent workforces trained on proprietary...SuggestedFull timeWork at officeVisa sponsorshipRelocation package- ...A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python...Suggested
- ...AI Infrastructure Engineer Spellbrush, the world's leading generative AI studio behind nijijourney, is looking for an AI Infrastructure Engineer to join us in building out end-to-end ML infrastructure to run our models on all platforms. What You'll Do Design...SuggestedWork at officeVisa sponsorship
$190k - $270k
...AI Chopping Block, Inc. is seeking an AI Infrastructure Engineer in San Francisco, CA, to ensure the optimal operation of user-facing services and production systems. This role requires expertise in building infrastructure with Ansible, Terraform, and Kubernetes, along...Suggested$175k - $300k
...This role sits at the intersection of platform engineering, site reliability, and applied ML systems. The... ...reliability, scalability, and operability of Meshy's AI model serving stack, along with core engineering infrastructure. The team operates a conventional production...SuggestedWork at officeRemote workFlexible hours- ...Cloudflare seeks a Senior Systems Engineer to help build and operate the AI Gateway at the network edge. You’ll move from half‑formed ideas to production‑ready systems across distributed services, high‑throughput APIs, and developer workflows, collaborating with Workers...
$180k - $225k
AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both...Full timeImmediate startRemote work$156.06k - $211.14k
Afresh, the AI platform for grocery, began by tackling the most complex problem in... ...feed them is not.As a Senior AI Platform Engineer, you build the AI and data platform that... ...top of it, and the evaluation and serving infrastructure underneath. Your "customers" are Afresh'...Full timeLive inWork at officeLocal areaRemote workWork from homeHome officeFlexible hours3 days per week$151.8k - $265.35k
The OpportunityAdobe empowers individuals and organizations to create exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity across design, imaging, motion, and personalization.We are seeking...Full timeTemporary workLocal areaWorldwide$269.1k - $307.2k
Distinguished AI Engineer (Agentic AI Platform) At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine...Full timePart timeWork at officeLocal area$192k - $259.8k
...stories, and career news. Job Summary: Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance... ...ll work closely with our agent developers, product engineers, and an embedded SRE partner, sitting at the intersection...Full timeWork at officeImmediate startWorldwideMonday to FridayFlexible hours$200k - $230k
...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco... ...are seeking a Director of AI Platform Engineering to lead the design, development, and... ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent...Ongoing contractFull timeCasual workWork at officeFlexible hours- ...the limits of what's possible.As a Lead Software Engineer at JP Morgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an... ...software applications and systems.Collaborate with AI teams to translate computational requirements...
- ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part of an... ...deploy secure, scalable cloud platforms optimized for AI/ML workloads.Partner with AI teams to translate...For contractors
$310k - $400k
...The "API-First World" graphic novel to understand the bigger picture and our vision at Postman.The OpportunityAs the Head of AI Platform Engineering at Postman, you will lead the alignment of AI development with our growing API platform. You will drive the AI roadmap with...Work at officeFlexible hours3 days per week- A leading AI research firm in San Francisco seeks a Staff Infrastructure Engineer to identify and resolve infrastructure bottlenecks and design large-scale systems for AI training. The ideal candidate has over 3 years of experience in infrastructure engineering and strong...
$220k - $300k
...Front to run their customer operations. AI is reshaping what's possible in this space... ...a group of our strongest applied AI and engineering talent whose work underpins every AI-... ...Building and maintaining the foundational AI infrastructure (LLM integrations, RAG pipelines,...Work at officeImmediate startRemote workWork from homeMonday to Friday- Grow Therapy in San Francisco is seeking a Senior AI Enablement Engineer to define how AI transforms operations across the organization. You will design and build foundational AI infrastructure that enhances efficiency. Responsibilities include implementing AI systems,...Flexible hours3 days per week
$300k - $330k
...by our unique combination of proprietary infrastructure and software, we empower over 250,000... ...you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster... ...a Director, Data and Knowledge Platform Engineering (based in San Francisco) to own the architecture...Temporary workLocal areaWorldwide$229.9k - $262.4k
Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area$229.9k - $262.4k
Senior Lead AI Engineer (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area- ...Overview Join a boutique quantamental hedge fund as our Lead AI Platform Engineer. Spearhead the buildout of a new internal data lake and platform. Build AI/ML-powered systems and hire, onboard, and manage a small team. Disrupt a mature industry and redefine how our strategies...Full timeImmediate start
- ...Senior Lead Software Engineer Be an integral part of an agile team that's constantly... ...JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part... ...scalable cloud platforms optimized for AI/ML workloads. Partner with AI teams...For contractors
$197.3k - $225.1k
Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area- ...About Zed Zed is building the first AI-native, licensed neobank in the Philippines... ...positioned to solve this. We are Stanford engineers and former YC founders who have spent... ...We're hiring an engineer to own the infrastructure layer behind our production AI agents....Full time
$171k - $240k
...and control spend effortlessly. Brex's AI-native automation and world-class service... ...grow your career. AI at Brex AI Engineering at Brex is redefining how businesses run... ...financial data with product and platform infrastructure, we're turning complex financial...Full timeWork at officeRemote workWork from homeShift work$150k - $210k
AI Engineer - Agentic Automation Location: Remote Compensation: $150,000 - $210,000 Join a rapidly growing company disrupting the... ...work at the intersection of AI, product development, and infrastructure — quickly turning ideas into working prototypes and scaling...Full timeLocal areaImmediate startRemote workFlexible hours$180k - $400k
About the Role We're a pre-seed AI-powered HR tech startup based in San Francisco,... ...domains. We're looking for a mid-level AI Engineer (2-8 years of experience) who is... ...generation (RAG) pipelines and retrieval infrastructure — including vector databases, embeddings...Full timeRelocation- ...TL;DR: If you: have a demonstrated track record of turning ambitious AI ideas into products people actually use; move effortlessly between research and engineering; have shipped something extraordinary, at work or outside of it; are relentlessly curious...Full timeRelocation package
- ...Job Description Job Description Benefits: ~401(k) AI Infrastructure Engineer / MLOps San Francisco Bay Area, CA (100% Onsite) EITACIES is looking for an experienced AI Infrastructure Engineer to support large scale AI and machine learning platforms running...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- senior infrastructure engineer San Francisco, CA
- infrastructure engineering manager San Francisco, CA


