AI Principal Engineer
Accellor
Job Description
Job Description
Accellor is an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities. Our mission is to operationalize AI at scale and unlock sustained enterprise value.
Our offerings span AI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. By leveraging design thinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability.
With a proven track record of enabling Fortune 100 enterprises and global innovators, Accellor stands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear: to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation.
Technical Architect — AI Systems & Platform Internals
Experience: 10–12 Years
Role Type: Technical Architect / Staff-Level Systems Architect
Role Summary
Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads.
This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.
The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.
Key Responsibilities :
1. AI Systems Architecture
Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.
Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.
Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.
2. Inference Runtime & Model Serving
Architect high-throughput, low-latency inference systems across large-scale GPU clusters.
Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.
Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.
Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.
3. GPU, Kernel & Distributed Performance
Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.
Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.
Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.
4. Context Engineering
Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.
Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.
Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.
5. Cost Optimization Frameworks
Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.
Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.
Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.
Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.
6. Training & Research Infrastructure
Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.
Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity.
Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution.
7. Release Safety, Validation & Evaluation Gates
Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.
Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.
Ensure platform optimizations do not reduce safety, grounding, quality, or user trust.
8. Reliability, Observability & Production Operations
Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.
Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops.
Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.
Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure.
9. Agentic & Multimodal Platform Internals
Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.
Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints.
Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads.
10. Technical Leadership
Work closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.
Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions.
Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.
Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.
Requirements
Required Qualifications:
- 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
- Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript .
- Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
- Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
- Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
- Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
- Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
- Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
- Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.
Preferred Qualifications:
- Experience working on LLM inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms.
- Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving.
- Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or LLM cost dashboards.
- Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems.
- Experience with large-scale distributed training, RL infrastructure, checkpointing, ML compiler optimizations, model graph transformations, or training runtime systems.
- Experience designing release gates, regression detection systems, canary systems, CI/CD validation frameworks, and production safety controls for performance-sensitive infrastructure.
- Experience with evals, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring.
Technical Skill Areas:
AI Systems: LLM serving, inference runtime, training infrastructure, post-training workflows, agent systems, multimodal models
Inference: batching, routing, KV-cache, streaming, latency optimization, model serving, tensor parallelism, pipeline parallelism
Performance Engineering: CUDA, Triton, GPU profiling, kernel optimization, memory bandwidth, communication libraries, distributed execution
Context Engineering: prompt architecture, dynamic context assembly, RAG, memory, context compression, context ranking, source grounding, permission-aware retrieval
Cost Optimization: token budgeting, caching, model routing, fallback strategies, cost telemetry, batching, async workflows, cost-quality trade-offs
Distributed Systems: scheduling, orchestration, reliability, fault tolerance, observability, scalability, service design
ML Frameworks: PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Ray
Infrastructure: Kubernetes, Docker, Terraform, CI/CD, cloud platforms, Linux systems, networking, storage
Safety & Validation: evals, release gates, canaries, regression testing, model behavior validation, rollout safety
Candidate Profile:
The ideal candidate is a senior hands-on architect who can operate across the full AI systems stack.
They should be able to discuss GPU memory bottlenecks, distributed inference, model-serving reliability, context quality, cost optimization, release validation, eval pipelines, observability, and production rollout with engineering teams, while also explaining architecture decisions clearly to senior leadership.
The candidate should not be limited to architecture diagrams. They must be capable of reviewing implementation quality, identifying bottlenecks, debugging production issues, challenging weak assumptions, and converting repeated failures into stronger platform abstractions.
This role requires the judgment of a senior architect, the debugging mindset of a systems engineer, and the ownership mindset required for production AI infrastructure.
$220k - $350k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.PRINCIPAL AI ENGINEER, SPECIAL PROGRAMSThis team focuses on engineering and deploying AI capabilities (models, APIs, tools, and integrations) for...SuggestedPermanent employmentTemporary workLocal areaImmediate startWeekend work$160k - $220k
...Fortinet, our mission is to safeguard people, devices, and data everywhere.Fortinet is seeking an experienced and innovative Principal AI Security Engineer to join our Corporate Information Security team. As an AI Security Engineer, you will play a crucial role in ensuring...SuggestedFull timeWork experience placementWorldwide$296.3k - $423.9k
...on a global scale. We are looking for a Principal Technical Lead Manager (TLM) to lead the... ...Trajectory Generation team within the Embodied AI organization. This role combines deep... ...You will lead a high-performing team of engineers building ML-driven trajectory generation...SuggestedFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$139.9k - $274.8k
A leading technology company is seeking a Principal Design Engineer to join the AI Silicon Engineering team. This role focuses on the design and development of high-performance digital logic for cloud infrastructure. Candidates must have a strong background in digital...Suggested$204k - $337k
...either Sunnyvale, CA or New York, NY. Team Overview: The Video AI team sits at the heart of our LinkedIn’s ambitious growth strategy... ...open and safe community for all.Responsibilities: As a Senior Engineering Manager you will lead a team of 15-20 engineers and technical...SuggestedFor contractorsWork at officeFlexible hours- Locations available: San Diego and San Jose, California or Austin, TexasNXP is searching for a hands-on AI Compiler Engineer who thrives at the convergence of cutting-edge AI, compiler tech, and hardware design. Here, you’ll not only architect and scale a production-class...Full timeWork at officeLocal area
$177k - $226k
...critical care through our rapid seizure detection technology, come join the movement!Position Overview:The Senior Manager, Applied AI Engineering is a senior individual contributor role with broad ownership across Ceribell's internal AI engineering portfolio. This person...For contractorsWork at officeLocal areaImmediate startRemote work$249k
...careers and have the flexibility, benefits, and support to do their best work. Join us and build for travelers everywhere.Principal Data & AI Engineer, Reporting and InsightsIntroduction to the Team: Our Technology team partners across Expedia Group to create innovative...Full timeWork at office$270k - $340k
...supports and celebrates all of our team members.What You’ll Do:As a Principal AI and ML fundamentalist who is an expert at developing cutting-... ..., and deployment. Collaborate with other researcher engineers to prototype and validate complex solutions from academic literature...Local area$148.7k - $297.3k
...for as well as a best place to work for diversity, working mothers, female executives, and scientists.THE OPPORTUNITYThis Principal AI/ML Engineer position can work out of our Santa Clara, CA location.The Principal ML Ops Engineer will lead the technical execution of Abbott...$250.44k - $375.67k
...processes, and foster a friendly, rewarding, and diverse environment for every OK-er.About the OpportunityWe are looking for a Principal AI Engineer to lead the architecture and deployment of large-scale, LLM-powered conversational Chatbot systems serving both enterprise...$118.8k - $190k
...Candidate Account, please Sign-In before you apply.Job Description:Broadcom’s GTO team is seeking a high-energy, visionary Principal AI Development Engineer to serve as a strategic leader and the primary architect of our AI transformation. In this Level 5 role, you are not...Full timeLocal area$190.2k - $360.5k
The Opportunity We are looking for a Principal AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product and platform capabilities. This role sits at the intersection of large-scale systems engineering, applied AI, and production...Full timeTemporary workLocal areaRemote workWorldwide$204k - $337k
...realize their greatest potential.Title and SummaryPrincipal AI Platform Engineer - AI Center of ExcellenceWho is Mastercard?Mastercard is a... ...for all.Overview:The AI Center of Excellence is seeking a Principal AI Infrastructure Engineer to build and scale next-generation...Full timePart timeWorldwideFlexible hours$123.24k - $200k
...Overview of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for Business Intelligence Innovation (AI4BII) Center, you will join an exciting global team dedicated to generating crucial business intelligence insights that shape TSMC's strategic...Full timeWork at office$184k
...Sr. Director, Search AI - Product EngineeringWe exist to wow our customers. We know we’re doing the right thing when we hear our... ...influence company-wide metrics, and lead a world-class team across engineering and machine learning.What You Will DoOwn and evolve Coupang’s...Temporary workFlexible hours- ...ServiceNow in Santa Clara, CA seeks a hands-on Cyber Defense Engineering architect to tackle hard, undefined security problems for a platform... ...used by thousands of enterprises. You’ll craft threat models, AI security controls, and governance, partnering with the product...
$313.06k
...friendly, rewarding, and diverse environment for every OK-er. About The Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment to spearhead the design, development, and optimization of intelligent agent...$25k
...We are looking for a Marketing AI Engineer to sit at the intersection of marketing strategy and cutting-edge AI technology. This 4-month contract role is responsible for accelerating adoption of AI within the marketing team by building, managing, and optimizing AI-powered...Remote jobFull timeContract workTemporary workWork at officeWorldwide- ...Company Description We're an Industrial AI start-up founded by a Stanford professor and led by recognized leaders in Data Intelligence... ...support where GenAI relies on proprietary Explainable AI engines for SME-explainable insights from historical data. Job...Remote jobFull timeContract workPart timeFor contractorsWork at office
$200k - $400k
...decided by those who field intelligent machines at scale. At Scout AI, we’re developing Fury, the first robotic foundation model for... ...work. The Role We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration Team with a deep passion for developing...Full timeRelocation package$180k
...About the Position As a Reasoning engineer, you will build frameworks to improve the reasoning capability, build distributed reinforcement learning systems, techniques for inference time compute (e.g. tree search and planning), and develop environments for agents....Full timeRelocation$130k - $220k
...About Eudia: Eudia is redefining the future of legal work with AI-powered Augmented Intelligence, enabling Fortune 500 legal teams... ...: We are looking to hire a solutions-minded, AI-focused engineer to join our growing team in Palo Alto. In this role, the...Full time$50k - $120k
...Who are we? Mission Altimate AI, founded in 2022 in San Francisco, is revolutionizing enterprise data operations through... ..., we're positioned at the forefront of the AI-powered data engineering revolution. You can read more about us in a recently published...Full timeWorldwide- ...world conflict. This mission will ask everything of us: urgency, precision, and relentless work. The Role We're looking for an AI Engineer to join the Fury Team with a deep passion for building next-generation autonomous systems. You’ll work across the stack of...Full timeRelocation package
- ...Intuition Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon... ...to accommodate family commitments. Meet our software engineers! Meet some of our software engineers who are shaping the future...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
- ..., runtime, and hardware teams on co-designed, high-performance AI applications. Integrate advances in model architecture, quantization... ...Bachelor’s or higher degree in computer science, electrical engineering, or a related field such as applied mathematics, physics, or...Full timeTemporary workFlexible hours
$313.06k
...A leading crypto exchange is seeking a Principal AI Engineer to architect and deploy large-scale, LLM-powered Chatbot systems for both enterprise and consumer use. This hands-on role requires defining technical vision and delivering production-grade platforms. Ideal candidates...- ...The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate... ...the role We are seeking a talented and driven ML performance engineer to optimize and scale state-of-the-art foundation models on...Full timeTemporary workLocal areaFlexible hours
- Yoh is seeking a Principal Technical Support Engineer to lead customer deployments and platform bring-up for next‑gen AI networking. You will optimize performance on high-speed Ethernet fabrics and GPU clusters, working across SONiC-based systems, switch ASICs, and related...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Principal Engineer. Be the first to apply!
- ai ml engineer Mountain View, CA
- machine learning ai engineer Mountain View, CA
- ai developer Mountain View, CA
- ai engineer Mountain View, CA
- ai prompt engineer Mountain View, CA
- senior ai engineer Mountain View, CA
- senior principal engineer Mountain View, CA
- data center chief engineer Mountain View, CA
- chief engineer Mountain View, CA
- senior chief engineer Mountain View, CA



