AI Systems engineer
Tilted AI
Systems Engineer, Training Infrastructure Tilted · San Francisco, CA [(hybrid) / Remote-US] · Full-time · Founding team · 2 openings About Tilted Tilted is an agent economy. Not a chat box — an operating system where every unit of work is a typed, owned, logged job on a NATS job mesh, where agents hold wallets, call each other as services, and settle payments in real time. Job state is reconstructed from an event log, not held in process memory, which is why a job survives a backend restart and any node can read any node's work. If that architecture appeals to you, the rest of this posting will too. We're standing up an in-house pretraining program — staged, starting at a few hundred GPUs and climbing from there over the next few years. You're one of the first two infrastructure engineers on it, working directly with the Founding Pretraining Lead. They own the recipe. You own the machine that makes it. What you'll own The training stack. Distributed training across data, tensor, pipeline, and expert parallelism. Standing it up, scaling it, and keeping it correct when the topology changes underneath you. MFU. Model FLOPs utilization is a first-class tracked metric here and you are the person accountable for it. At our later phases, one point of MFU is millions of dollars — this is not a vanity number. Fault tolerance. Checkpoint and restart at scale, node-failure detection and automatic recovery, silent data corruption detection, loss-spike rollback. Runs are measured in weeks; nodes die every day. The run should not care. The MoE communication problem. Expert-parallel all-to-all is the bottleneck that defines throughput for sparse models. Owning it is most of the job at later phases. The data pipeline at scale. Tokenization, sharding, streaming loaders that keep thousands of GPUs fed without stalling. Mesh integration. Wiring the whole pipeline — ingest, filter, train, eval — as typed jobs on TiltedOS's praxis mesh, so it schedules, retries, and audits itself. Observability. The dashboards and alerts that tell us a run is diverging before we've burned a week of compute finding out. First 90 days Stand up the training monorepo inside TiltedOS with the pretraining lead. Bring up the first reserved cluster and get a real multi-node run training end to end. Ship checkpoint/restart that survives a deliberately killed node without human intervention. Get the ablation grid launching as scheduled mesh jobs, with results flowing back automatically. What we're looking for Required You've run distributed training on multi-node GPU clusters and debugged it when it broke. Any framework, any scale — what matters is that you've been the person who fixed it. Strong PyTorch internals: autograd, memory behavior, profiling, and where the time actually goes. Real experience with NCCL and collective communication — and the ability to tell a network problem from a kernel problem from a data-loader problem. Systems fundamentals: Linux, networking, storage, containers. You can read a flame graph and a nsys trace. Comfort with the on-call reality of long runs. Something will break at 3am during a two-week job, and you'll want to know why rather than just restart it. Nice to have CUDA/Triton kernel work, or custom kernel integration. FSDP, DeepSpeed, Megatron-LM, TorchTitan, or equivalent at four-figure GPU counts and above. MoE-specific systems work: expert parallelism, all-to-all optimization, capacity/dropping tradeoffs. FP8 / FP4 mixed-precision training in production. High-performance storage and checkpointing at petabyte scale. Experience with NATS, event-sourced systems, or job meshes — you'll be building on one. How you work Instrument first, guess second. You treat a 5% throughput regression as a bug, not a rounding error. You'd rather automate the recovery than be the recovery. Why this role is different Not a cog role. Two engineers, one ML lead. You'll touch every layer of the stack from the kernel to the scheduler, and your name is on the throughput number. Infrastructure that's already interesting. You're not bolting a training pipeline onto generic cloud primitives. You're building it inside an OS where every job is already typed, owned, audited, and metered — and making a frontier training program run as jobs on that mesh is a genuinely novel systems problem. A roadmap with real scale at the end. We start at hundreds of GPUs, staged and honest. The plan doesn't stop there, and neither does the difficulty. Compensation [Base band] · [Meaningful founding equity] · [Benefits] · [Hardware + conference budget] We publish our bands. Tilted is an equal opportunity employer. #J-18808-Ljbffr
$190.2k - $360.5k
The Opportunity We are looking for a Principal AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product and platform capabilities. This role sits at the intersection of large-scale systems engineering, applied AI, and production...SuggestedFull timeTemporary workLocal areaRemote workWorldwide- ...Technical Staff with expertise in generative modelling to work in a collaborative, hybrid setting. You’ll develop and deploy advanced AI models for biotech and pharmaceutical applications. Qualified candidates will have significant experience in machine learning and be...Suggested
- ...AI Engineer, Search & Knowledge Systems About Pi Pi is building an agentic product security platform for teams that need to secure software at the speed they build it. Modern development is accelerating, but security knowledge is still scattered across code, tickets...SuggestedFull time
- ...Known - Conversational AI Engineer, System Prompt and Orchestration ~ San Francisco, CA (In-Person) ~170k-220k Cash + Equity Known is a matchmaker that talks to users and supports them like a friend. Our mission is to empower humanity by applying general intelligence...SuggestedFull time
$139k - $257.55k
...management, brand consistency, reusable design systems, and collaboration workflows that empower... ...team is exploring the next generation of AI-native creative systems that redefine how... ....We are looking for forward-thinking engineers who are excited to explore ambiguous...SuggestedFull timeTemporary workLocal areaWorldwide$150k - $230k
About the Role This is a founding engineer role at a small, high-impact startup building the evidence infrastructure layer for safety-critical AI — the system that proves a model works, and keeps proving it across its entire lifecycle. The company is focused on diagnostic...Full timeVisa sponsorshipShift work$190k - $286k
Amplitude is the leading AI analytics platform, helping over 4,700 customers—including... ...the Role We're looking for a Senior AI Engineer to help advance Amplitude's AI products... ...AI creates the most leverage, build the systems that deliver it, and establish the patterns...Full timeHome officeFlexible hours$175k - $225k
...AI/LLM Systems Software Engineer Location: San Francisco, CA Work Type: In-person Employment: Full-time Experience Required: 2+ years Salary Range: $175,000 - $225,000 per year Visa Sponsorship: Not available The Role As a founding engineer...Full time$149k - $240k
...Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP... ...re assembling a diverse, world-class team—engineers, designers, researchers, and product minds... ...best practices for distributed AI systems. Work closely with AI researchers, infrastructure...Full timeTemporary workLocal areaFlexible hours- ...last five years alone. Learn more at bishopfox.com or follow us on social media. Who You Are The Agentic AI Software Engineer – Cybersecurity Systems designs, develops, and deploys advanced AI-driven software solutions to enhance cybersecurity detection, response...Local areaWork from home
- ...About unitQ At unitQ , we leverage AI and advanced analytics to enable businesses... ...Opportunity We’re looking for an AI Software Engineer to lead features that use AI and... ...stakeholders to design, build, and monitor robust systems for data ingestion, cleaning, and...Flexible hours
$124k - $280k
...SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies... ...on developing and implementing advanced AI and ML solutions to drive innovation and... ...and optimising algorithms, models, and systems to enable intelligent decision-making and...Full timeH1b- ...within you. Job Description: Lead Software Engineer (Data technologies) Location: Sandy, UT... ...skilled Builder to design and implement AI-driven solutions that scale across cloud... ...Prototype and productionize LLM-based systems for dynamic SQL generation, document search...
$175k - $215k
...RoleMost companies are experimenting with AI in their GTM motion. We are not... ...exists to change that. You will build the systems that tie those signals together, eliminate... ...best practices, MCP architecture, prompt engineering standards, and enablementWrite documentation...Full time- THE GLOBAL LEADER IN DATA & ANALYTICS RECRUITMENTHarnham Search and Selection Company Number: 05723485Harnham Search and Selection is a registered company in England and Wales. Reg no. 05723485Harnham Europe Limited Company Number: 09956940Harnham GmbH HRB: 196954Harnham...Work at office
- Job Description:This is a fullstack engineering role with a strong frontend emphasis. You’ll have ownership across the product stack, from users interfaces to infrastructure supporting AI models. Location: New York, NY or Boston, MA or San Francisco, CA (3 days/week)Salary...3 days per week
$149k - $240k
Who We AreHP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global... ...assembling a diverse, world-class team—engineers, designers, researchers, and product... ...privacy best practices for distributed AI systems.Work closely with AI researchers, infrastructure...Full timeTemporary workLocal areaFlexible hours- ...San Francisco office.As part of Schwab’s AI Strategy & Transformation team (AI.x), you... ...company. AI.x brings together product, engineering, strategy, and risk specialists to set enterprise... ..., and operations. You’ll ensure that the systems you build are robust, reliable, and well-...Full timeWork at office
$160k - $200k
...the ER becomes their first touchpoint with the healthcare system—driving over $300B in avoidable costs every year.By... ...The Role:We're looking for an entrepreneurial Software Engineer with interest in agentic AI and at least 2-3 years of general engineering experience...Temporary workWork at officeMonday to FridayMonday to Thursday$171k - $240k
..., and control spend effortlessly. Brex’s AI-native automation and world-class service... ...you need to grow your career.AI at BrexAI Engineering at Brex is redefining how businesses run... ...finances by building intelligent, autonomous systems directly into the Brex platform. Our...Work at officeRemote workWork from homeShift work$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world’... ...work spans research into self-constructing systems, the development of the most reliable... ...of 20+ F500s. What We Are Looking ForAI Engineers build and operate production AI systems that...Work at office3 days per week$128k - $252.5k
...span from account executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking... ...health and clinical trials, autonomous systems and edge AI, and renewable energy. Key responsibilities...Local areaVisa sponsorship- We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits... ....Benchling is building Intelligence Engineering & Enablement, a small autonomous team within... ...partnership with our Data, Analytics & Systems team. We span the bridge between departmental...Work at officeLocal areaRemote workRelocationRelocation packageFlexible hours3 days per week
$120k - $200k
..., and you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster. Here... ...edge AI tooling across the spectrum: from prompt engineering and in-context learning to fine-tuned models and agentic systems, choosing the right approach for each problem....Temporary workLocal areaWorldwide- ...teams bring product management, design, and engineering together with the autonomy to move... ...engineering role. You Are You are an exceptional AI-native software engineer who has built... ...fundamentals with product judgment, systems thinking, and a high bar for quality. You...Permanent employmentFull timeWork experience placementLive inWork at officeLocal area
$115k - $175k
Gong harnesses the power of AI to transform how revenue teams win. The Gong Revenue AI Operating System unifies data, insights, and workflows into a single, trusted system... ...with high executive visibility.As a GTM AI Engineer, you'll design, build, and ship AI-powered infrastructure...Remote workWork from homeFlexible hours$163k - $246.5k
...Semgrep gets smarter as you build, with AI that learns your context to cut false positives... ...at semgrep.dev.About the roleAs an AI engineer, you’ll apply LLM technologies throughout... ...compensate every Semgrep employee with a system that equally rewards those who are vocal...Currently hiringWork at officeLocal areaRemote workWeekend work3 days per week$110.7k - $372.9k
...fifty million Americans rely on a healthcare system whose decision-making has become slow,... ...the care they need. Deloitte has a new AI-first effort, backed by $1B in committed... ...months, not into a lab. As an Agentic AI Engineer, you will design, build, and operationalize...Local areaVisa sponsorship$150k - $210k
AI Engineer - Agentic AutomationLocation: RemoteCompensation: $150,000 - $210,000Join a rapidly growing company disrupting the trucking... ...driven applications. This role focuses on building agentic AI systems and automation tools that integrate with our ecosystem and deliver...Local areaImmediate start$186.5k - $328.5k
...world's leading enterprises orchestrate AI-powered work. Our vision is to expand human... ...of work with AI. About the roleAs an AI engineer at WRITER, you'll be at the forefront of... ...building and deploying machine learning or AI systems in a production environment.Proficiency...Full timeWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Systems engineer. Be the first to apply!
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- systems engineer San Francisco, CA
- system engineer contract San Francisco, CA





