Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Principal Engineer

Accellor

Job Description

Job Description

Accellor is an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities. Our mission is to operationalize AI at scale and unlock sustained enterprise value.

Our offerings span AI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. By leveraging design thinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability.

With a proven track record of enabling Fortune 100 enterprises and global innovators, Accellor stands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear: to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation.

Technical Architect — AI Systems & Platform Internals

Experience: 10–12 Years
Role Type: Technical Architect / Staff-Level Systems Architect

Role Summary

Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads.

This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.

The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.

Key Responsibilities :

1. AI Systems Architecture

Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.

Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.

Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.

2. Inference Runtime & Model Serving

Architect high-throughput, low-latency inference systems across large-scale GPU clusters.

Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.

Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.

Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.

3. GPU, Kernel & Distributed Performance

Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.

Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.

Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.

4. Context Engineering

Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.

Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.

Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.

5. Cost Optimization Frameworks

Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.

Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.

Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.

Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.

6. Training & Research Infrastructure

Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.

Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity.

Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution.

7. Release Safety, Validation & Evaluation Gates

Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.

Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.

Ensure platform optimizations do not reduce safety, grounding, quality, or user trust.

8. Reliability, Observability & Production Operations

Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.

Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops.

Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.

Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure.

9. Agentic & Multimodal Platform Internals

Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.

Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints.

Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads.

10. Technical Leadership

Work closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.

Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions.

Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.

Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.

Requirements

Required Qualifications:

  • 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript .
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
  • Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.

Preferred Qualifications:

  • Experience working on LLM inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms.
  • Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving.
  • Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or LLM cost dashboards.
  • Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems.
  • Experience with large-scale distributed training, RL infrastructure, checkpointing, ML compiler optimizations, model graph transformations, or training runtime systems.
  • Experience designing release gates, regression detection systems, canary systems, CI/CD validation frameworks, and production safety controls for performance-sensitive infrastructure.
  • Experience with evals, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring.

Technical Skill Areas:

AI Systems: LLM serving, inference runtime, training infrastructure, post-training workflows, agent systems, multimodal models

Inference: batching, routing, KV-cache, streaming, latency optimization, model serving, tensor parallelism, pipeline parallelism

Performance Engineering: CUDA, Triton, GPU profiling, kernel optimization, memory bandwidth, communication libraries, distributed execution

Context Engineering: prompt architecture, dynamic context assembly, RAG, memory, context compression, context ranking, source grounding, permission-aware retrieval

Cost Optimization: token budgeting, caching, model routing, fallback strategies, cost telemetry, batching, async workflows, cost-quality trade-offs

Distributed Systems: scheduling, orchestration, reliability, fault tolerance, observability, scalability, service design

ML Frameworks: PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Ray

Infrastructure: Kubernetes, Docker, Terraform, CI/CD, cloud platforms, Linux systems, networking, storage

Safety & Validation: evals, release gates, canaries, regression testing, model behavior validation, rollout safety

Candidate Profile:

The ideal candidate is a senior hands-on architect who can operate across the full AI systems stack.

They should be able to discuss GPU memory bottlenecks, distributed inference, model-serving reliability, context quality, cost optimization, release validation, eval pipelines, observability, and production rollout with engineering teams, while also explaining architecture decisions clearly to senior leadership.

The candidate should not be limited to architecture diagrams. They must be capable of reviewing implementation quality, identifying bottlenecks, debugging production issues, challenging weak assumptions, and converting repeated failures into stronger platform abstractions.

This role requires the judgment of a senior architect, the debugging mindset of a systems engineer, and the ownership mindset required for production AI infrastructure.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Principal Engineer in San Francisco, CA vacancy
  •  ...Join SignalFire’s Talent Network for Principal AI/ML Engineer Roles at VC-Backed Startups This is not an application for a specific job. Instead, this is a way to get on the radar of VC-backed startups that are actively hiring AI/ML talent. If you have any questions,... 
    Suggested
    Full time

    Signal Fire Inc

    San Francisco, CA
    20 hours ago
  • $300k - $400k

     ...WHO WE ARE  Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (...  ...world. To learn more, go to . Role Description   As a Principal AI/ML Engineer in our AdTech team, you will be a key individual contributor... 
    Suggested
    Full time

    Zeta Global

    San Francisco, CA
    20 hours ago
  • $308k - $423.5k

     ...we power the shop local movement. If you believe in community, come join ours. About this role: We are seeking a  Principal ML / AI Engineer to be a  company-level technical thought leader and practitioner to help shape the future of Data and AI at Faire. This... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire Faire

    San Francisco, CA
    20 hours ago
  • $229k - $281k

    MissionAs a AI Accelerated Engineering Lead, you will guide teams in the shift toward AI-accelerated engineering and define clear guardrails for...  ..., we are hiring at these locations: San Francisco* Senior Principal: $229,000-$281,000Washington DC* Senior Principal: $210,00... 
    Suggested
    Temporary work
    Local area
    Shift work

    Slalom

    San Francisco, CA
    4 days ago
  •  ...build and manage their workforce through an intelligent, auditable AI platform that spans the entire employee lifecycle. As part of a confidential search, Scovai is seeking a Principal Generative AI Engineer to serve as the technical cornerstone of a rapidly growing AI... 
    Suggested
    Full time

    CONFIDENTIAL Scovai

    San Francisco, CA
    8 days ago
  •  ...Overview We are seeking a Principal GenAI Architect / Forward Deployed Principal Engineer to lead the design, delivery, and productionization of enterprise-grade Generative AI solutions within a highly regulated banking environment. This role will partner directly with... 
    Temporary work

    Ontrac Solutions

    San Francisco, CA
    a month ago
  • $190.2k - $360.5k

    The Opportunity We are looking for a Principal AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product and platform capabilities. This role sits at the intersection of large-scale systems engineering, applied AI, and production... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • Innovaccer seeks a Principal AI Engineer to build production-grade AI systems at scale. You will take ideas from paper to prototype to production, design multi-model pipelines, and ensure product-level accuracy through rigorous evaluation. Role requires deep expertise in... 

    Socket

    San Francisco, CA
    4 days ago
  • $260k - $275k

    Medium is seeking a Senior Principal Software Engineer in San Francisco to lead the design and implementation of AI security solutions. This role requires over 15 years in software engineering, with expert skills in Java, Spring, and cloud platforms such as AWS and Azure... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $73.5k - $212.28k

     ...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design...  ...- Work with cross-functional teams to incorporate AI into various applications- Drive initiatives that enhance project... 
    Full time
    H1b

    PwC

    San Francisco, CA
    1 day ago
  • $73.5k - $212.28k

     ...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design...  ...requirements.The OpportunityAs part of the People Tech & AI team you will lead the design, build, and operation of scalable... 
    Full time
    Work experience placement
    H1b
    Remote work

    PwC

    San Francisco, CA
    4 days ago
  • $161.7k - $303.3k

     ...organization is rebuilding how marketing teams operate — not by layering AI tools on top of existing workflows, but by replacing them. The...  ...people doing it.You'll manage a team of 6-8 Forward-Deployed AI Engineers embedded across GMI's paid media, lifecycle marketing, data... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • $293.5k

    General Information Job Title Expert Senior Manager, AI Engineering Job ID 104335 Work Areas Analytics, Data & Research, Management Consulting, Technology & Engineering Employment Type Permanent Full-Time Location(s) Atlanta, Austin, Boston... 
    Permanent employment
    Full time
    Apprenticeship
    Work at office
    Local area
    Work from home
    Home office
    3 days per week

    Bain & Company

    San Francisco, CA
    4 days ago
  • $144k - $240k

    Lila Sciences is seeking a Sr Principal / Principal Software Engineer to join their innovative team in San Francisco, CA. You will design and build AI-driven applications, focusing on performance, reliability, and cross-functional collaboration with scientists. Ideal candidates... 
    Flexible hours

    jobr.pro

    San Francisco, CA
    3 days ago
  • $228k - $340k

     ...Everlaw is looking for a Staff/Principal AI Engineer to help design and execute our overall technical strategy of applying AI in our product, directly contributing to our company’s vision of being the AI leader in legal discovery and litigation technology. AI is central... 
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours
    3 days per week

    Everlaw

    Oakland, CA
    20 hours ago
  •  ...Your Role The AI & Machine Learning team works in partnership across the enterprise to accelerate business outcomes by applying...  ...Reporting to the Director, AI & Machine Learning, the Data Scientist, Principal will lead the development and deployment of novel applications... 
    Full time
    Part time
    Work at office
    Local area
    Work from home
    Home office
    2 days per week

    Socket.dev

    Oakland, CA
    19 hours ago
  • $310k - $400k

     ...The "API-First World" graphic novel to understand the bigger picture and our vision at Postman.The OpportunityAs the Head of AI Platform Engineering at Postman, you will lead the alignment of AI development with our growing API platform. You will drive the AI roadmap with... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    3 days ago
  • $200k - $230k

     ...for expertise, scale, and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco, CA / Boston, MA /...  ...enterprise big data platform.We are seeking a Director of AI Platform Engineering to lead the design, development, and scaling of our enterprise... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Flexible hours

    SS&C Technologies

    San Francisco, CA
    20 hours ago
  • About the teamThe Applied AI Engineering (AAE) team is responsible for helping developers and enterprises turn the potential of generative AI into real-world impact. We act as trusted advisors and technical partners to customers and ecosystem partners, helping identify... 
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    4 days ago
  • $220k - $300k

     ...,000 companies rely on Front to run their customer operations. AI is reshaping what's possible in this space, and Front is building...  ...AI Platform team - a group of our strongest applied AI and engineering talent whose work underpins every AI-powered experience across... 
    Work at office
    Immediate start
    Remote work
    Work from home
    Monday to Friday

    FrontApp

    San Francisco, CA
    3 days ago
  • Aware Health is hiring a Director of Engineering in San Francisco for a full-time, on-site/ hybrid schedule (3 days in SF). You will lead and scale our engineering team, shaping AI-driven healthcare solutions and platforms to advance orthopedic care. Ideal candidates bring... 
    Full time

    ApplyMint

    San Francisco, CA
    2 days ago
  • $156.4k - $301k

     ...better working world. The Opportunity As an Associate Director in EY’s Forward Deployed Engineering team, you will support the design, development, and deployment of AI-driven, data-centric solutions within strategic client environments. This role blends strong technical... 
    Summer holiday
    Local area
    Flexible hours

    EY

    San Francisco, CA
    1 day ago
  •  ...AI has changed software development, but security hasn't caught up — until now. Corridor is redefining product security for the AI...  ...on open models in government and academia. We're hiring an AI Engineer, Product to make the AI systems that power Corridor's product measurably... 
    Full time

    Corridor

    San Francisco, CA
    20 hours ago
  •  ...Parasail is redefining AI infrastructure by enabling seamless deployment across a distributed network of GPUs, optimizing for cost...  ...mission. About the Role We’re looking for a hungry, creative engineer who thrives in a high-trust, high-velocity environment. You’ll... 
    Full time

    Parasail

    San Francisco, CA
    20 hours ago
  • $150k - $350k

     ...About Collate   Collate is an AI document generation platform for life sciences. We automate paperwork with AI, helping our customers...  ...at Y Combinator and founder of Lever. Our AI researchers, engineers, and designers have worked at Google, Nvidia, Meta, Netflix, Amazon... 
    Full time

    Collate

    San Francisco, CA
    20 hours ago
  •  ...We’re hiring an AI Engineer to build the intelligence layer for the leading AI companion for language learning. You’ll own the core AI systems powering personalized conversations for millions of users. Language learning shouldn’t feel like an app — it should feel like... 
    Full time

    Pingo پینگو

    San Francisco, CA
    20 hours ago
  •  ...About the role We’re seeking an experienced engineer to deploy enterprise-grade AI solutions, focusing on Retrieval-Augmented Generation (RAG) pipelines and large language model (LLM) workflows. This role is vital to expanding our reach with Fortune 500 and enterprise... 
    Full time

    StackAI

    San Francisco, CA
    20 hours ago
  • Chime is seeking a Director of Engineering for the AI & App Experience (AAX) organization to lead engineering strategy for Jade, the AI-powered financial assistant, and the AI platform. You will guide multi-quarter plans, build a leadership bench, and drive architectural... 

    Chime

    San Francisco, CA
    3 days ago
  • $150k - $250k

     ...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect...  ...the board-members of 20+ F500s. What We Are Looking For AI Engineers build and operate production AI systems that deliver business value... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Distyl Ai

    San Francisco, CA
    20 hours ago
  •  ...Role: ML/AI Engineers (This role is open to US Citizens, Green Card holders, GC-EAD only. We do not sponsor visas.)   Summary: Adidev is looking for an adept Machine Learning Engineer to take the helm in deploying advanced machine learning models, with a special... 
    Full time
    Remote work
    Visa sponsorship
    Relocation package

    Adidev Technologies Inc

    San Francisco, CA
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Principal Engineer. Be the first to apply!