AI Systems engineer
Tilted AI
Systems Engineer, Training Infrastructure Tilted · San Francisco, CA [(hybrid) / Remote-US] · Full-time · Founding team · 2 openings About Tilted Tilted is an agent economy. Not a chat box — an operating system where every unit of work is a typed, owned, logged job on a NATS job mesh, where agents hold wallets, call each other as services, and settle payments in real time. Job state is reconstructed from an event log, not held in process memory, which is why a job survives a backend restart and any node can read any node's work. If that architecture appeals to you, the rest of this posting will too. We're standing up an in-house pretraining program — staged, starting at a few hundred GPUs and climbing from there over the next few years. You're one of the first two infrastructure engineers on it, working directly with the Founding Pretraining Lead. They own the recipe. You own the machine that makes it. What you'll own The training stack. Distributed training across data, tensor, pipeline, and expert parallelism. Standing it up, scaling it, and keeping it correct when the topology changes underneath you. MFU. Model FLOPs utilization is a first-class tracked metric here and you are the person accountable for it. At our later phases, one point of MFU is millions of dollars — this is not a vanity number. Fault tolerance. Checkpoint and restart at scale, node-failure detection and automatic recovery, silent data corruption detection, loss-spike rollback. Runs are measured in weeks; nodes die every day. The run should not care. The MoE communication problem. Expert-parallel all-to-all is the bottleneck that defines throughput for sparse models. Owning it is most of the job at later phases. The data pipeline at scale. Tokenization, sharding, streaming loaders that keep thousands of GPUs fed without stalling. Mesh integration. Wiring the whole pipeline — ingest, filter, train, eval — as typed jobs on TiltedOS's praxis mesh, so it schedules, retries, and audits itself. Observability. The dashboards and alerts that tell us a run is diverging before we've burned a week of compute finding out. First 90 days Stand up the training monorepo inside TiltedOS with the pretraining lead. Bring up the first reserved cluster and get a real multi-node run training end to end. Ship checkpoint/restart that survives a deliberately killed node without human intervention. Get the ablation grid launching as scheduled mesh jobs, with results flowing back automatically. What we're looking for Required You've run distributed training on multi-node GPU clusters and debugged it when it broke. Any framework, any scale — what matters is that you've been the person who fixed it. Strong PyTorch internals: autograd, memory behavior, profiling, and where the time actually goes. Real experience with NCCL and collective communication — and the ability to tell a network problem from a kernel problem from a data-loader problem. Systems fundamentals: Linux, networking, storage, containers. You can read a flame graph and a nsys trace. Comfort with the on-call reality of long runs. Something will break at 3am during a two-week job, and you'll want to know why rather than just restart it. Nice to have CUDA/Triton kernel work, or custom kernel integration. FSDP, DeepSpeed, Megatron-LM, TorchTitan, or equivalent at four-figure GPU counts and above. MoE-specific systems work: expert parallelism, all-to-all optimization, capacity/dropping tradeoffs. FP8 / FP4 mixed-precision training in production. High-performance storage and checkpointing at petabyte scale. Experience with NATS, event-sourced systems, or job meshes — you'll be building on one. How you work Instrument first, guess second. You treat a 5% throughput regression as a bug, not a rounding error. You'd rather automate the recovery than be the recovery. Why this role is different Not a cog role. Two engineers, one ML lead. You'll touch every layer of the stack from the kernel to the scheduler, and your name is on the throughput number. Infrastructure that's already interesting. You're not bolting a training pipeline onto generic cloud primitives. You're building it inside an OS where every job is already typed, owned, audited, and metered — and making a frontier training program run as jobs on that mesh is a genuinely novel systems problem. A roadmap with real scale at the end. We start at hundreds of GPUs, staged and honest. The plan doesn't stop there, and neither does the difficulty. Compensation [Base band] · [Meaningful founding equity] · [Benefits] · [Hardware + conference budget] We publish our bands. Tilted is an equal opportunity employer. #J-18808-Ljbffr Tilted AI
- ...Technical Staff with expertise in generative modelling to work in a collaborative, hybrid setting. You’ll develop and deploy advanced AI models for biotech and pharmaceutical applications. Qualified candidates will have significant experience in machine learning and be...Suggested
- ...Known - Conversational AI Engineer, System Prompt and Orchestration ~ San Francisco, CA (In-Person) ~170k-220k Cash + Equity Known is a matchmaker that talks to users and supports them like a friend. Our mission is to empower humanity by applying general intelligence...SuggestedFull time
$120k - $180k
About the Role This is a founding engineering role at a seed-stage construction AI company building the intelligence layer that transforms messy blueprints... ...have outsized influence over the technical direction, system architecture, and culture of the team. This is a full...SuggestedFull timeWork at officeRemote workRelocation$180k - $400k
About the Role We're a pre-seed AI-powered HR tech startup based in San Francisco, building... ..., product, and design to ship agentic systems that automate complex, multi-step... ...domains. We're looking for a mid-level AI Engineer (2-8 years of experience) who is comfortable...SuggestedFull timeRelocation$225k - $255k
About the Role This is a founding-level AI engineering role at an early-stage B2B SaaS pricing intelligence startup based in San Francisco... ...sits at the intersection of LLM infrastructure, evaluation systems, and revenue-critical product outcomes. You'll build the feedback...SuggestedFull timeRelocation$90k - $200k
About the Role A fast-growing, Y Combinator-backed B2B SaaS startup in the sales automation space is looking for an AI/LLM Engineer to join their team in Munich. The company automates quote and order processing for distributors and manufacturers — helping sales teams eliminate...Full timeVisa sponsorship$120k - $180k
About the Role This is a founding engineering role at a seed-stage AI startup building the intelligence layer for the construction industry. The company... ..., and multimodal AI to prototype and ship production systems on real-world, noisy construction documents. The focus...Full timeWork at officeRelocation$225k - $255k
...Role We're a small, product-focused team building an AI-powered B2B pricing platform that helps companies continuously... ...guidance and discount governance. As our Founding AI Engineer, you'll own the evaluation systems, feedback loops, and LLM infrastructure that allow our...Full timeRelocation$150k - $200k
...founding team in San Francisco building an AI-first platform that organises real-world... ...-to-end by AI agents. As a Founding AI Engineer, you'll own two of the most technically challenging... ...filtering, graph embeddings, and ranking systems to match people into highly compatible...Full timeRemote workVisa sponsorship- ...demonstrated track record of turning ambitious AI ideas into products people actually use;... ...move effortlessly between research and engineering; have shipped something extraordinary,... ...closer. You'll work on the fundamental systems that allow Archie to reason, learn, and...Full timeRelocation package
$200k - $400k
...Overview Job Title: Founding AI Engineer Location: San Francisco, CA (On-site; relocation assistance available) Compensation: $200,000–... ...AI Engineer to help design, build, and launch a multi-agent AI system that will redefine the future of talent evaluation. You’ll join...Full timeImmediate startRelocationVisa sponsorshipRelocation packageFlexible hours$229.9k - $262.4k
Senior Lead AI Engineer(MLX, Agentic AI, Gen AI platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time...Full timePart timeLocal area$150k - $250k
About the Role Join a Series A AI-powered engineering automation startup that is reimagining computer-aided engineering (CAE) for the AI era.... ...optimizing model accuracy to analyzing outputs and addressing system-level bottlenecks. Guide AI/ML projects end-to-end, from...Full timeWork at officeVisa sponsorship- ...people interact with the web by building AI agents that can reliably do everyday digital... ...web‑agent Work closely with product engineers to translate cutting‑edge AI capabilities... ...MFU, throughput, latency) Low level systems experience (Triton, CUDA) High IQ, high...Work at officeRelocationVisa sponsorship
- Do you want to work at the intersection of AI, product, engineering, and real -world operations? \ \ Join a fast -growing technology startup building intelligent systems that optimize workflows, improve efficiency and implement AI for operations in industrial environments...Full timeImmediate start
- ...technology that works for them, not against them. Revic is an AI-native sales acceleration engine built to uplift sales professionals. We handle the... ..., and Head of Engineering to design and build the systems that translate real customer conversations and GTM data...Remote jobFull time
- ...partnered with a well-funded startup in the AI Healthtech space. They're building a... ...care advocate to navigate the healthcare system on their behalf. The role will work a hybrid... ...best practices, and build with the next engineer in mind. What you'll bring: You...Full timeWork at officeRemote work
- ...revolutionizing software development with AI-powered formal verification. We've... ...About the role Join our team as an AI Engineer and help us push the boundaries of what's... ...Expertise in optimizing machine learning systems, including general techniques and LLM-specific...Full timeContract work
- ...As a Backend/AI Engineer at Emanate, you'll work on the core infrastructure that powers our AI revenue engine for industrial materials... ...integrate cutting-edge LLMs, optimize data pipelines, and build the systems that make autonomous revenue workflows possible at scale....Full time
- ...Be one of the founding engineers at Nen, shaping the AI layer that powers automation across enterprise desktop environments at scale. The role Build and extend a multi-model agent loop across leading AI providers Benchmark models across cost, latency, and reliability...Full time
- ...About the Role Fieldguide is building AI agents for the most complex audit and advisory... ...top-tier investors. As a Senior AI Engineer, Quality , you will own the evaluation... ...source of truth for all of our agentic systems and audit workflows Build observability...Remote jobFull timeWork at officeWork from homeFlexible hours
- ...Github, Gmail, Notion, Salesforce, etc. We are a small team of engineers wrangling problems from context to search, that help us provide... ...people actually stop scrolling to watch it. You'll own how the AI community meets Composio. What you'll do? Build in public...Full timePart timeRelocationFlexible hours
- ...is a scalable, data science-first growth engine that gives B2C teams predictive clarity into... ...We're also co-building alongside leading AI companies. We're looking for an AI Engineer who can build production-grade AI systems end-to-end - from prototype to pipeline to...Full timeShift workNight shiftWeekend work
$150k - $250k
...Description Max AI – Stripe for Healthcare Max AI is the World’s first human-free, fully-autonomous medical billing AI agent... ...research at MIT and Caltech for over 10 years. And our Head of Engineering was one of the earliest engineers at Figma. AI Engineer...Full time$180k - $300k
...About The Role You'll own the core AI systems that power Gamma: the models, prompts, and pipelines behind text, image, and layout... ...latency, and cost across our AI stack. You'll work closely with engineering and product to ship improvements that millions of users feel...Full timeWork at officeImmediate startWork from home- ...Forward Deployed AI Engineer The opportunity We are looking for a Forward Deployed AI Engineer to serve as the critical bridge between... ...generative biology platform integrates seamlessly with their systems. You will own the full lifecycle of customer deployments -...Full timeShift work
- ...that helps people be exceptional and thrive at work through human, AI, and software-based coaching. We’re on a mission to provide... ...The Role We’re looking for a product-minded senior software engineer with experience working with data to help us build a differentiated...Remote jobFull timeHome office
- LiteLLM is the world's most popular AI Gateway, trusted by companies like Adobe, Netflix... ..., LLM APIs, FastAPI As the Backend LLM Engineer, you'll be responsible for ensuring... ...and CTO on critical projects including: System design for supporting provider level features...Full timeImmediate start
- ...For phase 1, we're training foundational AI models that create building code-compliant... ...models, geometric and physics engines from scratch to transform how the world designs... ...This is not a BIM modeling role. It is a systems and algorithms role focused on building the...Full timeVisa sponsorship
- ...created Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments of... ...try. Sign up today (it’s free)! Role Overview As an AI Engineer at Fathom, you'll be hands-on with LLMs, prototyping and...Full timeWork at officeRemote work3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Systems engineer. Be the first to apply!
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- systems engineer San Francisco, CA
- system engineer contract San Francisco, CA



