Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Principal Engineer

Accellor

Job Description

Job Description

Accellor is an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities. Our mission is to operationalize AI at scale and unlock sustained enterprise value. 

Our offerings span AI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. By leveraging design thinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability. 

With a proven track record of enabling Fortune 100 enterprises and global innovators, Accellor stands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear: to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation. 

Technical Architect — AI Systems & Platform Internals

Experience: 10–12 Years
Role Type: Technical Architect / Staff-Level Systems Architect

Role Summary

Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads.

This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.

The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.

Key Responsibilities :

1. AI Systems Architecture

Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.

Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.

Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.

2. Inference Runtime & Model Serving

Architect high-throughput, low-latency inference systems across large-scale GPU clusters.

Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.

Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.

Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.

3. GPU, Kernel & Distributed Performance

Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.

Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.

Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.

4. Context Engineering

Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.

Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.

Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.

5. Cost Optimization Frameworks

Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.

Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.

Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.

Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.

6. Training & Research Infrastructure

Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.

Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity.

Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution.

7. Release Safety, Validation & Evaluation Gates

Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.

Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.

Ensure platform optimizations do not reduce safety, grounding, quality, or user trust.

8. Reliability, Observability & Production Operations

Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.

Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops.

Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.

Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure.

9. Agentic & Multimodal Platform Internals

Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.

Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints.

Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads.

10. Technical Leadership

Work closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.

Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions.

Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.

Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.

Requirements

Required Qualifications:

  • 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript .
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
  • Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.

Preferred Qualifications:

  • Experience working on LLM inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms.
  • Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving.
  • Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or LLM cost dashboards.
  • Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems.
  • Experience with large-scale distributed training, RL infrastructure, checkpointing, ML compiler optimizations, model graph transformations, or training runtime systems.
  • Experience designing release gates, regression detection systems, canary systems, CI/CD validation frameworks, and production safety controls for performance-sensitive infrastructure.
  • Experience with evals, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring.

Technical Skill Areas:

AI Systems: LLM serving, inference runtime, training infrastructure, post-training workflows, agent systems, multimodal models

Inference: batching, routing, KV-cache, streaming, latency optimization, model serving, tensor parallelism, pipeline parallelism

Performance Engineering: CUDA, Triton, GPU profiling, kernel optimization, memory bandwidth, communication libraries, distributed execution

Context Engineering: prompt architecture, dynamic context assembly, RAG, memory, context compression, context ranking, source grounding, permission-aware retrieval

Cost Optimization: token budgeting, caching, model routing, fallback strategies, cost telemetry, batching, async workflows, cost-quality trade-offs

Distributed Systems: scheduling, orchestration, reliability, fault tolerance, observability, scalability, service design

ML Frameworks: PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Ray

Infrastructure: Kubernetes, Docker, Terraform, CI/CD, cloud platforms, Linux systems, networking, storage

Safety & Validation: evals, release gates, canaries, regression testing, model behavior validation, rollout safety

Candidate Profile:

The ideal candidate is a senior hands-on architect who can operate across the full AI systems stack.

They should be able to discuss GPU memory bottlenecks, distributed inference, model-serving reliability, context quality, cost optimization, release validation, eval pipelines, observability, and production rollout with engineering teams, while also explaining architecture decisions clearly to senior leadership.

The candidate should not be limited to architecture diagrams. They must be capable of reviewing implementation quality, identifying bottlenecks, debugging production issues, challenging weak assumptions, and converting repeated failures into stronger platform abstractions.

This role requires the judgment of a senior architect, the debugging mindset of a systems engineer, and the ownership mindset required for production AI infrastructure.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the AI Principal Engineer in San Francisco, CA vacancy
  • DescriptionAbout the RoleAs a Principal AI Engineer at Innovaccer, you will be one of our most senior engineers and you will stay deep in the code. You'll personally design and build the hardest, highest-impact parts of our platform — backend services, data pipelines, and... 
    Suggested
    Full time
    Temporary work

    Innovaccer

    San Francisco, CA
    5 days ago
  •  ...We are hiring a Applied AI Engineer for a leading fintech company's AI Team to serve as the team's technical anchor in the United States. You will lead two primary streams: AI Coding tooling and developer experience - bringing the engineering depth of Silicon Valley... 
    Suggested

    ByLabs

    San Francisco, CA
    5 days ago
  • $175k - $200k

     ...H2O.ai is on a mission to democratize AI for Good. As the world’s leading agentic AI company, H2O.ai converges Generative and...  ...information, visit About This Opportunity We are looking for a Principal AI Engineer who builds things that matter. You will design and ship end-... 
    Suggested
    Full time
    Remote work
    Worldwide
    Flexible hours

    h2o.ai

    San Francisco, CA
    16 days ago
  •  ...build and manage their workforce through an intelligent, auditable AI platform that spans the entire employee lifecycle. As part of a confidential search, Scovai is seeking a Principal Generative AI Engineer to serve as the technical cornerstone of a rapidly growing AI... 
    Suggested
    Full time

    CONFIDENTIAL Scovai

    San Francisco, CA
    9 days ago
  •  ...capital and venture buyout strategies, each powered by an AI operating system and team of leading technologists, entrepreneurs...  ...We're hiring an ambitious, self-initiated, product-minded Principal Applied AI Engineer to build the agentic workflows that power Redesign’s most... 
    Suggested

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $290k

     ...data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a...  ...in all we do. About the Role Nscale is looking for a Principal AI Engineer (Specialised) to lead the inference and post-training pillar... 
    Flexible hours

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...At Atlassian, we are seeking a Senior Principal Forward Deployed Engineer (FDE) to advance AI-powered solutions for customers. You will work with product, engineering and design teams to deliver end-to-end deployments and real-world impact. You’ll lead across the stack... 

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $302k - $335k

    About the TeamThe Applied AI Engineering team is responsible for helping customers turn frontier AI capabilities into real products, workflows, and business impact. We act as trusted technical partners across solution design, architecture, implementation, evaluation, and... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $135.2k - $198.3k

     ...your future. Levi's is accelerating the next phase of its SAP and AI transformation by embedding intelligent automation and AI into...  ...processes that power the company. As Senior Manager of SAP AI Engineering, you will lead the deployment and scale-up of AI-enabled ERP... 
    Full time
    Work at office
    Remote work
    3 days per week

    Levi Strauss & Co

    San Francisco, CA
    3 days ago
  • $293.6k - $335.1k

     ...Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has...  ...to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough... 
    Full time
    Part time
    Local area

    SupportFinity

    San Francisco, CA
    2 days ago
  • $302k - $335k

    About the TeamThe Applied AI Engineering team partners closely with customers to help them turn frontier AI capabilities into real products, workflows, and business impact. We act as trusted technical partners across strategy, solution design, architecture, implementation... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  • $212k - $265k

     ...and unlock incredible career growth opportunities, join us, and build real world value.Join Ripple Labs Inc. as the Manager, AI Platform Engineering in Chicago, IL, and spearhead our bold venture in AI platform development! This remarkable opportunity enables you to craft... 
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    5 days ago
  •  ...Capital One is seeking a Director, AI Engineer (Remote Eligible) to lead AI product delivery and platform initiatives. You will oversee end-to-end AI software lifecycle, mentor leaders, and drive responsible AI practices across engineering teams. You will collaborate... 
    Remote job

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $244.7k - $279.2k

     ...At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been...  ...committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough... 
    Remote job
    Full time
    Part time
    Local area

    Jobleads-US

    San Francisco, CA
    5 days ago
  •  ...Capital One is seeking an experienced Director of AI Engineering to lead the Intelligent Foundations and Experiences team. You will oversee foundation model work, LLM inference, guardrails, and observability, while guiding a cross-functional team to deliver production... 
    Remote job

    Jobleads-US

    San Francisco, CA
    5 days ago
  •  ...Structure Therapeutics Inc. seeks a Director of AI Engineering to build and scale production AI capabilities across drug discovery, clinical development, manufacturing, and business operations. The role is hands‑on and player‑coach, with 60–70% direct technical contribution... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $148.1k - $282.1k

    The Opportunity We're looking for a Principal Product Manager with deep fluency in the AI ecosystem to lead the evolution of our engineering platform. The rapid adoption of AI is fundamentally reshaping how we build, ship, and scale software at Adobe. To keep pace, we... 
    Full time
    Temporary work
    Local area
    Immediate start
    Worldwide

    Adobe Systems

    San Francisco, CA
    1 day ago
  • $180k - $220k

     ...most personal data people share, trust and safety at global scale, AI features shipping into consumer products, and regulation that...  ...traditional legal team role. It's for a tech-forward builder, an engineer or product person, who can find where AI can do legal work that... 
    Full time
    Contract work
    Work experience placement
    Immediate start
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Match Group

    San Francisco, CA
    1 day ago
  •  ...Job Description Job Description About the Role As a Principal AI Engineer at Innovaccer, you will be one of our most senior engineers and you will stay deep in the code. You'll personally design and build the hardest, highest-impact parts of our platform — backend... 
    Temporary work

    Innovaccer Analytics

    San Francisco, CA
    8 days ago
  •  ...Okta seeks a seasoned Engineering Manager for Governance Intelligence to lead ML and software engineering efforts, shaping AI-powered identity governance solutions. You will guide architecture, development, and deployment across cloud and on‑prem environments, ensuring... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • Job ID: 29357116Reference Number: 26-00976Title: Python Backend Engineer / AI EngineerPosted Date: 2026-09-24Contact: Vishika ChaudharyContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Python Backend Engineer / AI EngineerJob... 
    Work at office

    HAN Staffing

    San Francisco, CA
    2 days ago
  • $171k - $240k

     ...gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual...  ...resources, and support you need to grow your career.AI at BrexAI Engineering at Brex is redefining how businesses run their finances by building... 
    Work at office
    Remote work
    Work from home

    Brex

    San Francisco, CA
    5 days ago
  • $175k - $215k

     ...relevant. About the RoleMost companies are experimenting with AI in their GTM motion. We are not experimenting. We are building...  ...across the Revenue Org: best practices, MCP architecture, prompt engineering standards, and enablementWrite documentation and run enablement... 
    Full time

    Samba TV

    San Francisco, CA
    1 day ago
  • $122k - $240.5k

    Position Summary Agentic AI is moving from experimentation to production, and organizations everywhere are racing to figure...  ...how to build it responsibly and at scale. We're growing a team of engineers who want to work at the center of that shift: designing and... 
    Work at office
    Local area
    Visa sponsorship
    Shift work

    Deloitte

    San Francisco, CA
    1 day ago
  •  ...with solutions created by the #1 company in e-signature and contract lifecycle management (CLM). What you'll do As an AI Agentic Engineer on the Platform AI Engineering team, you will drive the next wave of automation in IT operations by designing and deploying agentic... 
    Permanent employment
    Full time
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    San Francisco, CA
    5 days ago
  • THE GLOBAL LEADER IN DATA & ANALYTICS RECRUITMENTHarnham Search and Selection Company Number: 05723485Harnham Search and Selection is a registered company in England and Wales. Reg no. 05723485Harnham Europe Limited Company Number: 09956940Harnham GmbH HRB: 196954Harnham...
    Work at office

    Harnham

    San Francisco, CA
    1 day ago
  •  ...promised to change everything across a business—until generative AI. Today, AI is the number one driver of business reinvention. And...  ...and help lead that change.You Are:As a Snowflake Advanced AI Engineer, you will design, build, and operationalize artificial intelligence... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Shift work

    Accenture

    San Francisco, CA
    3 days ago
  • $120k - $200k

    About AirwallexAirwallex is the AI-native financial operating system for a real-time, intelligent economy. More than 676,000 businesses...  ...cutting-edge AI tooling across the spectrum: from prompt engineering and in-context learning to fine-tuned models and agentic systems... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    5 days ago
  • $149k - $240k

    Who We AreHP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies...  ..., and collaborates.We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating... 
    Full time
    Temporary work
    Local area
    Flexible hours

    HP IQ

    San Francisco, CA
    5 days ago
  • $127k - $175k

    AI Software Engineer - HP IQDescription -Who We AreHP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates.We’re assembling a diverse, world... 
    Full time
    Temporary work
    Local area
    Relocation
    Flexible hours
    Shift work

    Juniper Networks

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Principal Engineer. Be the first to apply!