SRE: AI Inference Platform & ML Systems
Cohere
A leading AI technology company in San Francisco is looking for a Site Reliability Engineer to build high-performance, scalable systems for model serving. You will collaborate with teams to deploy optimized NLP models and ensure high availability. Ideal candidates have 5+ years of experience, particularly in Kubernetes and large-scale infrastructure. The role is focused on automating services and maintaining system reliability in a fast-paced environment. The company values diversity and provides competitive benefits including 6 weeks of vacation. #J-18808-Ljbffr Cohere
- Senior Software Engineer - AI Inference Systems AI needs far more compute than exists today, but simply... ...runtime performance. Shape both the platform and the engineering culture as the team... .... Experience building or operating ML inference or model serving systems. A...Platform
$167.2k - $209k
...DigitalOcean is expanding its AI Infrastructure layer to... ...2 to join our AI Inference Data Plane team. In... ...intersection of distributed systems and specialized AI... ...SLOs to ensure superior platform health. What You’ll Bring... ...as code. AI/ML Domain Knowledge: Hands...PlatformLocal areaRemote workWorldwideFlexible hours$172.5k - $210k
Check out 30 new AI Systems Engineer opportunities posted on AI Chopping... ...job entails building shared platform capabilities that unblock... ...backend, data systems, and applied ML engineering domains.... ...system reliability, real-time inference observability, sovereign data...PlatformLocal area$165k - $330k
...Baseten powers mission-critical inference for dynamic AI companies, enabling models... ...applications on Baseten’s platform. You’ll own the journey with... ...Develop and maintain software systems and product features using... ...Python due to its relevance in ML projects. Drive customer...PlatformWork experience placementFlexible hours- AI Systems Engineer - Codex Core Agents The Codex Core Agents team builds... ...harness, model interaction, inference, sandboxed execution,... ...across low‑level systems and ML workflows, able to debug Codex... ...sandboxing, virtualization, cloud platforms, or ML systems. Enjoy...Platform
- AI Systems Engineer - Codex Core Agents Location San Francisco Employment Type Full time... ...harness issues, model behavior, inference/runtime issues, and product failures.... ...tooling, sandboxing, virtualization, cloud platforms, or ML systems. Enjoy working across layers:...PlatformFull timeWork at officeLocal areaRelocation packageFlexible hours
$300k
...building a next-generation AI and cloud platform designed for startups... ...model training and inference, with flexible... ...AI workloads, automate systems at petascale, and be part... .... Collaborate with ML, networking, and platform... ...years of experience in SRE, DevOps, or...PlatformPermanent employmentFlexible hours- ...seeking a Technical Program Manager for Inference in San Francisco, California. This role is... ...coordinating strategic initiatives across inference systems, ensuring reliability and smooth... ...in technical program management within ML/AI systems, strong skills in stakeholder communication...
$180k - $250k
...who can help scale the platform. This is a strong fit for... ..., and distributed systems. Role 1: Software Engineer... ...systems that power secure AI agent execution. This person... ...Engineer The SRE role is focused on keeping... ...with AI infrastructure, ML workloads, GPU clusters,...PlatformFull timeImmediate start$175k - $200k
Staff Platform Engineer - Agentic AI Systems, IFS The Loops Staff Platform Engineer - Agentic AI Systems, IFS The... ...Collaborate closely with infrastructure, SRE, and product teams to ensure platform... ...ecosystem. Experience with AI or ML-driven orchestration or agentic frameworks...PlatformFull timeFor contractorsFlexible hours- ...first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental... ...from the underlying hardware. Our platform intelligently partitions workloads... ...of Technical Staff focused on ML systems and inference. In this role, you will design and...Platform
- Luma AI in San Francisco is building cutting-edge multimodal AI systems. We seek a seasoned Platform Engineer to ship new model architectures into our inference engine and optimize deployments across clusters. You will collaborate with research, engineering and infra, develop...Platform
- ...Software Engineers focused on AI Infrastructure to build the systems that enable frontier... ...orchestration, large-scale inference systems, performance optimization, and developer platforms that allow applied scientists... ...with GPU-based ML workloads or distributed training...PlatformInternshipImmediate start
- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This... ...designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team, you will...PlatformWorldwide
- Capital One is seeking an experienced AI Engineer to develop and optimize AI-powered products... ...should have substantial experience in AI/ML, and programming skills in Python, Go,... ...will design software components for AI systems and lead projects that contribute to Capital...Platform
- Gravity Engineering Services Pvt Ltd. in San Francisco is looking for a specialized engineer to advance the efficiency of ML inference systems. The role encompasses algorithm design, system optimization, and the integration of RL-driven training techniques. Ideal candidates...
$215.2k - $245.6k
...leading financial services company seeks a Lead AI Engineer in San Francisco to develop and optimize AI systems. You will collaborate with engineers and researchers... ...expertise in AI algorithm development and cloud platforms. The salary range for this position is $215,200 -...Platform- MeshyAI is seeking a Platform Engineer to enhance our AI inference platform in San Francisco. You will design and develop core capabilities, focusing on resource management and service orchestration. The ideal candidate holds a relevant degree and has experience in backend...PlatformRemote jobFlexible hours
$192k - $240k
A leading AI solutions provider in San Francisco is seeking a Senior AI/ML Engineer to architect the core platform for synthetic data generation and agentic workflows. The ideal candidate... ...of experience in cloud-native software systems, possesses deep knowledge of AI/ML...PlatformWork at office- ...designing, building, and scaling production ML systems. Responsibilities include leading end-to-end ML processes, developing scalable ML platforms, and implementing MLOps best practices.... ...offers opportunities to drive impactful AI solutions and collaborate with cross-functional...Platform
- ...at the intersection of platform engineering, site reliability, and applied ML systems. The function owns the... ...operability of Meshy’s AI model serving stack, along... ...for the AI inference platform, including inference... ...experience in observability, SRE, capacity planning,...PlatformWork at officeRemote workFlexible hours
- Lightning AI in San Francisco or Seattle is seeking a Senior Application Security Engineer to secure our AI/ML platforms and inference services. You will work with platform, ML, and infrastructure teams to identify risks and implement secure architectures. The role emphasizes...Platform
$225k - $250k
...leading open‑source AI coding agent.... ...that supports our AI inference pipelines, backend... ...beyond just keeping systems running—you’ll design... ...backend engineers and ML teams. As a Senior... ...the future of our platform. Responsibilities:... ...in infrastructure, SRE, or platform engineering...Platform- ..., we’re creating a platform to help businesses... ...customer experiences with AI. We are primarily... ...to ensure our systems are highly available... ...the foundation of SRE practices at Sierra... ...collaborating across product, ML, and core... ...infrastructure — optimizing inference performance,...PlatformFull timeFlexible hours
- LuminX is transforming warehouse operations with AI-powered camera systems that monitor each pallet move in real time, reading labels... .... You’ll work across NVIDIA Jetson/Rockchip platforms, hardware integration, and ML deployment, with frequent travel to customer sites...Platform
- ...BASETEN Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor, Notion,... ...Conviction. Join us and help build the platform engineers turn to to ship AI... ...technical work is needed. REQUIREMENTS AI/ML background and the ability to...PlatformFlexible hours
- Gravity Engineering Services Pvt Ltd. is looking for a skilled ML Systems Engineer to drive research and development in Physical AI. You'll design and build platforms for scalable and efficient model serving, enhancing both research and production systems for autonomous...Platform
$250k
...Ready to architect AI infrastructure that powers... ...generation research and cloud platforms? Join a stealth-mode... ...building a serverless inference platform, beginning... ...distributed inference systems to maximise GPU utilisation... ...distributed systems (ML inference, HPC, or...PlatformPermanent employment- ...scalability across Sierra's AI-driven... ...teams to ensure our systems are highly available... ...Partnering with product and platform engineers to design... ...the foundation of SRE practices at Sierra... ...across product, ML, and core... ...infrastructure — optimizing inference performance,...PlatformFull timeFlexible hours
- About the job Software Engineer Data/AI/Intelligent Systems Cisco is a leading technology company focused... ...data pipelines and building analytics platforms to support machine learning initiatives... ...Preferred Hands-on experience with AI/ML Familiarity with major cloud platforms...PlatformApprenticeship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE: AI Inference Platform & ML Systems. Be the first to apply!
- director of digital platform San Francisco, CA
- digital platform specialist San Francisco, CA
- power platform San Francisco, CA
- platform manager San Francisco, CA
- platform product manager San Francisco, CA
- data engineer machine learning San Francisco, CA
- machine learning San Francisco, CA
- machine learning intern San Francisco, CA
- internship machine learning San Francisco, CA
- machine learning researcher San Francisco, CA


