Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems engineer

Tilted AI

Systems Engineer, Training Infrastructure Tilted · San Francisco, CA [(hybrid) / Remote-US] · Full-time · Founding team · 2 openings About Tilted Tilted is an agent economy. Not a chat box — an operating system where every unit of work is a typed, owned, logged job on a NATS job mesh, where agents hold wallets, call each other as services, and settle payments in real time. Job state is reconstructed from an event log, not held in process memory, which is why a job survives a backend restart and any node can read any node's work. If that architecture appeals to you, the rest of this posting will too. We're standing up an in-house pretraining program — staged, starting at a few hundred GPUs and climbing from there over the next few years. You're one of the first two infrastructure engineers on it, working directly with the Founding Pretraining Lead. They own the recipe. You own the machine that makes it. What you'll own The training stack. Distributed training across data, tensor, pipeline, and expert parallelism. Standing it up, scaling it, and keeping it correct when the topology changes underneath you. MFU. Model FLOPs utilization is a first-class tracked metric here and you are the person accountable for it. At our later phases, one point of MFU is millions of dollars — this is not a vanity number. Fault tolerance. Checkpoint and restart at scale, node-failure detection and automatic recovery, silent data corruption detection, loss-spike rollback. Runs are measured in weeks; nodes die every day. The run should not care. The MoE communication problem. Expert-parallel all-to-all is the bottleneck that defines throughput for sparse models. Owning it is most of the job at later phases. The data pipeline at scale. Tokenization, sharding, streaming loaders that keep thousands of GPUs fed without stalling. Mesh integration. Wiring the whole pipeline — ingest, filter, train, eval — as typed jobs on TiltedOS's praxis mesh, so it schedules, retries, and audits itself. Observability. The dashboards and alerts that tell us a run is diverging before we've burned a week of compute finding out. First 90 days Stand up the training monorepo inside TiltedOS with the pretraining lead. Bring up the first reserved cluster and get a real multi-node run training end to end. Ship checkpoint/restart that survives a deliberately killed node without human intervention. Get the ablation grid launching as scheduled mesh jobs, with results flowing back automatically. What we're looking for Required You've run distributed training on multi-node GPU clusters and debugged it when it broke. Any framework, any scale — what matters is that you've been the person who fixed it. Strong PyTorch internals: autograd, memory behavior, profiling, and where the time actually goes. Real experience with NCCL and collective communication — and the ability to tell a network problem from a kernel problem from a data-loader problem. Systems fundamentals: Linux, networking, storage, containers. You can read a flame graph and a nsys trace. Comfort with the on-call reality of long runs. Something will break at 3am during a two-week job, and you'll want to know why rather than just restart it. Nice to have CUDA/Triton kernel work, or custom kernel integration. FSDP, DeepSpeed, Megatron-LM, TorchTitan, or equivalent at four-figure GPU counts and above. MoE-specific systems work: expert parallelism, all-to-all optimization, capacity/dropping tradeoffs. FP8 / FP4 mixed-precision training in production. High-performance storage and checkpointing at petabyte scale. Experience with NATS, event-sourced systems, or job meshes — you'll be building on one. How you work Instrument first, guess second. You treat a 5% throughput regression as a bug, not a rounding error. You'd rather automate the recovery than be the recovery. Why this role is different Not a cog role. Two engineers, one ML lead. You'll touch every layer of the stack from the kernel to the scheduler, and your name is on the throughput number. Infrastructure that's already interesting. You're not bolting a training pipeline onto generic cloud primitives. You're building it inside an OS where every job is already typed, owned, audited, and metered — and making a frontier training program run as jobs on that mesh is a genuinely novel systems problem. A roadmap with real scale at the end. We start at hundreds of GPUs, staged and honest. The plan doesn't stop there, and neither does the difficulty. Compensation [Base band] · [Meaningful founding equity] · [Benefits] · [Hardware + conference budget] We publish our bands. Tilted is an equal opportunity employer. #J-18808-Ljbffr Tilted AI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Systems engineer in San Francisco, CA vacancy
  •  ...Technical Staff with expertise in generative modelling to work in a collaborative, hybrid setting. You’ll develop and deploy advanced AI models for biotech and pharmaceutical applications. Qualified candidates will have significant experience in machine learning and be... 
    Suggested

    Latent Labs

    San Francisco, CA
    3 days ago
  •  ...Known - Conversational AI Engineer, System Prompt and Orchestration ~ San Francisco, CA (In-Person) ~170k-220k Cash + Equity Known is a matchmaker that talks to users and supports them like a friend. Our mission is to empower humanity by applying general intelligence... 
    Suggested
    Full time

    Known

    San Francisco, CA
    21 hours ago
  • $120k - $180k

    About the Role This is a founding engineering role at a seed-stage construction AI company building the intelligence layer that transforms messy blueprints...  ...have outsized influence over the technical direction, system architecture, and culture of the team. This is a full... 
    Suggested
    Full time
    Work at office
    Remote work
    Relocation

    Clera

    San Francisco, CA
    3 days ago
  • $180k - $400k

    About the Role We're a pre-seed AI-powered HR tech startup based in San Francisco, building...  ..., product, and design to ship agentic systems that automate complex, multi-step...  ...domains. We're looking for a mid-level AI Engineer (2-8 years of experience) who is comfortable... 
    Suggested
    Full time
    Relocation

    Clera

    San Francisco, CA
    3 days ago
  • $225k - $255k

    About the Role This is a founding-level AI engineering role at an early-stage B2B SaaS pricing intelligence startup based in San Francisco...  ...sits at the intersection of LLM infrastructure, evaluation systems, and revenue-critical product outcomes. You'll build the feedback... 
    Suggested
    Full time
    Relocation

    Clera

    San Francisco, CA
    3 days ago
  • $90k - $200k

    About the Role A fast-growing, Y Combinator-backed B2B SaaS startup in the sales automation space is looking for an AI/LLM Engineer to join their team in Munich. The company automates quote and order processing for distributors and manufacturers — helping sales teams eliminate... 
    Full time
    Visa sponsorship

    Clera

    San Francisco, CA
    1 day ago
  • $120k - $180k

    About the Role This is a founding engineering role at a seed-stage AI startup building the intelligence layer for the construction industry. The company...  ..., and multimodal AI to prototype and ship production systems on real-world, noisy construction documents. The focus... 
    Full time
    Work at office
    Relocation

    Clera

    San Francisco, CA
    3 days ago
  • $225k - $255k

     ...Role We're a small, product-focused team building an AI-powered B2B pricing platform that helps companies continuously...  ...guidance and discount governance. As our Founding AI Engineer, you'll own the evaluation systems, feedback loops, and LLM infrastructure that allow our... 
    Full time
    Relocation

    Clera

    San Francisco, CA
    3 days ago
  • $150k - $200k

     ...founding team in San Francisco building an AI-first platform that organises real-world...  ...-to-end by AI agents. As a Founding AI Engineer, you'll own two of the most technically challenging...  ...filtering, graph embeddings, and ranking systems to match people into highly compatible... 
    Full time
    Remote work
    Visa sponsorship

    Clera

    San Francisco, CA
    3 days ago
  •  ...demonstrated track record of turning ambitious AI ideas into products people actually use;...  ...move effortlessly between research and engineering; have shipped something extraordinary,...  ...closer. You'll work on the fundamental systems that allow Archie to reason, learn, and... 
    Full time
    Relocation package

    P-1 AI

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...Overview Job Title: Founding AI Engineer Location: San Francisco, CA (On-site; relocation assistance available) Compensation: $200,000–...  ...AI Engineer to help design, build, and launch a multi-agent AI system that will redefine the future of talent evaluation. You’ll join... 
    Full time
    Immediate start
    Relocation
    Visa sponsorship
    Relocation package
    Flexible hours

    Heaven and Hire Recruiting

    San Francisco, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer(MLX, Agentic AI, Gen AI platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    1 day ago
  • $150k - $250k

    About the Role Join a Series A AI-powered engineering automation startup that is reimagining computer-aided engineering (CAE) for the AI era....  ...optimizing model accuracy to analyzing outputs and addressing system-level bottlenecks. Guide AI/ML projects end-to-end, from... 
    Full time
    Work at office
    Visa sponsorship

    Clera

    San Francisco, CA
    3 days ago
  •  ...people interact with the web by building AI agents that can reliably do everyday digital...  ...web‑agent Work closely with product engineers to translate cutting‑edge AI capabilities...  ...MFU, throughput, latency) Low level systems experience (Triton, CUDA) High IQ, high... 
    Work at office
    Relocation
    Visa sponsorship

    Yutori

    San Francisco, CA
    1 day ago
  • Do you want to work at the intersection of AI, product, engineering, and real -world operations? \ \ Join a fast -growing technology startup building intelligent systems that optimize workflows, improve efficiency and implement AI for operations in industrial environments... 
    Full time
    Immediate start

    Eth Juniors

    San Francisco, CA
    21 hours ago
  •  ...technology that works for them, not against them. Revic is an AI-native sales acceleration engine built to uplift sales professionals. We handle the...  ..., and Head of Engineering to design and build the systems that translate real customer conversations and GTM data... 
    Remote job
    Full time

    Revic

    San Francisco, CA
    21 hours ago
  •  ...partnered with a well-funded startup in the AI Healthtech space. They're building a...  ...care advocate to navigate the healthcare system on their behalf. The role will work a hybrid...  ...best practices, and build with the next engineer in mind. What you'll bring: You... 
    Full time
    Work at office
    Remote work

    Drh Search Llc

    San Francisco, CA
    21 hours ago
  •  ...revolutionizing software development with AI-powered formal verification. We've...  ...About the role Join our team as an AI Engineer and help us push the boundaries of what's...  ...Expertise in optimizing machine learning systems, including general techniques and LLM-specific... 
    Full time
    Contract work

    Logical Intelligence

    San Francisco, CA
    21 hours ago
  •  ...As a Backend/AI Engineer at Emanate, you'll work on the core infrastructure that powers our AI revenue engine for industrial materials...  ...integrate cutting-edge LLMs, optimize data pipelines, and build the systems that make autonomous revenue workflows possible at scale.... 
    Full time

    Emanate Inc.

    San Francisco, CA
    21 hours ago
  •  ...Be one of the founding engineers at Nen, shaping the AI layer that powers automation across enterprise desktop environments at scale. The role Build and extend a multi-model agent loop across leading AI providers Benchmark models across cost, latency, and reliability... 
    Full time

    Nen

    San Francisco, CA
    21 hours ago
  •  ...About the Role Fieldguide is building AI agents for the most complex audit and advisory...  ...top-tier investors. As a Senior AI Engineer, Quality , you will own the evaluation...  ...source of truth for all of our agentic systems and audit workflows Build observability... 
    Remote job
    Full time
    Work at office
    Work from home
    Flexible hours

    Field Guide Inc

    San Francisco, CA
    21 hours ago
  •  ...Github, Gmail, Notion, Salesforce, etc. We are a small team of engineers wrangling problems from context to search, that help us provide...  ...people actually stop scrolling to watch it. You'll own how the AI community meets Composio. What you'll do? Build in public... 
    Full time
    Part time
    Relocation
    Flexible hours

    Composio

    San Francisco, CA
    21 hours ago
  •  ...is a scalable, data science-first growth engine that gives B2C teams predictive clarity into...  ...We're also co-building alongside leading AI companies. We're looking for an AI Engineer who can build production-grade AI systems end-to-end - from prototype to pipeline to... 
    Full time
    Shift work
    Night shift
    Weekend work

    Hilbert's Ai

    San Francisco, CA
    21 hours ago
  • $150k - $250k

     ...Description Max AI – Stripe for Healthcare Max AI is the World’s first human-free, fully-autonomous medical billing AI agent...  ...research at MIT and Caltech for over 10 years. And our Head of Engineering was one of the earliest engineers at Figma. AI Engineer... 
    Full time

    Maxai

    San Francisco, CA
    21 hours ago
  • $180k - $300k

     ...About The Role You'll own the core AI systems that power Gamma: the models, prompts, and pipelines behind text, image, and layout...  ...latency, and cost across our AI stack. You'll work closely with engineering and product to ship improvements that millions of users feel... 
    Full time
    Work at office
    Immediate start
    Work from home

    Gamma

    San Francisco, CA
    21 hours ago
  •  ...Forward Deployed AI Engineer The opportunity We are looking for a Forward Deployed AI Engineer to serve as the critical bridge between...  ...generative biology platform integrates seamlessly with their systems. You will own the full lifecycle of customer deployments -... 
    Full time
    Shift work

    Latent Labs

    San Francisco, CA
    21 hours ago
  •  ...that helps people be exceptional and thrive at work through human, AI, and software-based coaching. We’re on a mission to provide...  ...The Role  We’re looking for a product-minded senior software engineer with experience working with data to help us build a differentiated... 
    Remote job
    Full time
    Home office

    株式会社mento

    San Francisco, CA
    21 hours ago
  • LiteLLM is the world's most popular AI Gateway, trusted by companies like Adobe, Netflix...  ..., LLM APIs, FastAPI As the Backend LLM Engineer, you'll be responsible for ensuring...  ...and CTO on critical projects including: System design for supporting provider level features... 
    Full time
    Immediate start

    Litellm

    San Francisco, CA
    21 hours ago
  •  ...For phase 1, we're training foundational AI models that create building code-compliant...  ...models, geometric and physics engines from scratch to transform how the world designs...  ...This is not a BIM modeling role. It is a systems and algorithms role focused on building the... 
    Full time
    Visa sponsorship

    Augrade Private Limited

    San Francisco, CA
    21 hours ago
  •  ...created Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments of...  ...try. Sign up today (it’s free)! Role Overview As an AI Engineer at Fathom, you'll be hands-on with LLMs, prototyping and... 
    Full time
    Work at office
    Remote work
    3 days per week

    Fathom

    San Francisco, CA
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems engineer. Be the first to apply!