Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems & Platform Internals - Technical Architect

Accellor

Job Description

Job Description

Accellor is an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities. Our mission is to operationalize AI at scale and unlock sustained enterprise value.

Our offerings span AI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. By leveraging design thinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability.

With a proven track record of enabling Fortune 100 enterprises and global innovators, Accellor stands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear: to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation.

Technical Architect — AI Systems & Platform Internals

Experience: 10–12 Years
Role Type: Technical Architect / Staff-Level Systems Architect

Role Summary

Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads.

This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.

The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.

Key Responsibilities :

1. AI Systems Architecture

Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.

Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.

Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.

2. Inference Runtime & Model Serving

Architect high-throughput, low-latency inference systems across large-scale GPU clusters.

Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.

Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.

Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.

3. GPU, Kernel & Distributed Performance

Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.

Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.

Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.

4. Context Engineering

Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.

Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.

Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.

5. Cost Optimization Frameworks

Design and build cost optimization frameworks for large-scale LLM and GenAI workloads.

Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.

Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.

Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.

6. Training & Research Infrastructure

Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.

Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity.

Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution.

7. Release Safety, Validation & Evaluation Gates

Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.

Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.

Ensure platform optimizations do not reduce safety, grounding, quality, or user trust.

8. Reliability, Observability & Production Operations

Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.

Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops.

Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.

Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure.

9. Agentic & Multimodal Platform Internals

Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.

Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints.

Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads.

10. Technical Leadership

Work closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.

Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions.

Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.

Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.

Requirements

Required Qualifications:

  • 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript .
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
  • Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.

Preferred Qualifications:

  • Experience working on LLM inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms.
  • Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving.
  • Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or LLM cost dashboards.
  • Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems.
  • Experience with large-scale distributed training, RL infrastructure, checkpointing, ML compiler optimizations, model graph transformations, or training runtime systems.
  • Experience designing release gates, regression detection systems, canary systems, CI/CD validation frameworks, and production safety controls for performance-sensitive infrastructure.
  • Experience with evals, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring.

Technical Skill Areas:

AI Systems: LLM serving, inference runtime, training infrastructure, post-training workflows, agent systems, multimodal models

Inference: batching, routing, KV-cache, streaming, latency optimization, model serving, tensor parallelism, pipeline parallelism

Performance Engineering: CUDA, Triton, GPU profiling, kernel optimization, memory bandwidth, communication libraries, distributed execution

Context Engineering: prompt architecture, dynamic context assembly, RAG, memory, context compression, context ranking, source grounding, permission-aware retrieval

Cost Optimization: token budgeting, caching, model routing, fallback strategies, cost telemetry, batching, async workflows, cost-quality trade-offs

Distributed Systems: scheduling, orchestration, reliability, fault tolerance, observability, scalability, service design

ML Frameworks: PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Ray

Infrastructure: Kubernetes, Docker, Terraform, CI/CD, cloud platforms, Linux systems, networking, storage

Safety & Validation: evals, release gates, canaries, regression testing, model behavior validation, rollout safety

Candidate Profile:

The ideal candidate is a senior hands-on architect who can operate across the full AI systems stack.

They should be able to discuss GPU memory bottlenecks, distributed inference, model-serving reliability, context quality, cost optimization, release validation, eval pipelines, observability, and production rollout with engineering teams, while also explaining architecture decisions clearly to senior leadership.

The candidate should not be limited to architecture diagrams. They must be capable of reviewing implementation quality, identifying bottlenecks, debugging production issues, challenging weak assumptions, and converting repeated failures into stronger platform abstractions.

This role requires the judgment of a senior architect, the debugging mindset of a systems engineer, and the ownership mindset required for production AI infrastructure.

Vacancy posted 17 days ago
Similar jobs that could be interesting for youBased on the AI Systems & Platform Internals - Technical Architect in San Francisco, CA vacancy
  • $148.5k - $237.6k

     ...Finance Functional Architect to help us scale....  ...future of Finance systems by leveraging modern...  ..., analytics, and AI-driven solutions...  ...Operations (D365 FO) platform. Conduct fit/gap...  ...Collaborate with the internal audit and finance...  ...collaborate with technical and non-technical... 
    Platform
    Work experience placement

    Axon

    San Francisco, CA
    3 days ago
  • $149k - $175k

     ...harnesses the power of AI to transform how...  ...Revenue AI Operating System unifies data, insights...  ...work of your career.The Technical Architect is a specialized individual...  ...who ensures the Gong platform is seamlessly integrated...  ....Develop and refine internal technical best practice... 
    Platform
    Remote work
    Work from home
    Flexible hours

    Gong.io

    San Francisco, CA
    3 days ago
  • $148.19k - $231.98k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ....The Salesforce Cloud Technical Architect is a pre-sales...  ..., distributed platforms, mobile, and analytics...  ...architecture, software/systems engineering, cloud computing...  ...customers, partners and to internal audiences.Baseline... 
    Platform
    Full time
    Work experience placement
    Remote work

    Salesforce

    San Francisco, CA
    4 days ago
  • $190k - $232k

     ...way they connect their internal systems and interact with...  ...partner with leading platforms—Salesforce, ServiceNow...  ...functions, or activate AI-powered innovation, we...  ...ERP system). MuleSoft Architects are responsible for setting the overall technical direction of solutions... 
    Platform
    Temporary work
    Local area

    Slalom

    San Francisco, CA
    4 days ago
  • Block, Inc. is seeking a Legal Systems Engineer in San Francisco to design and evolve scalable legal tech ecosystems. You will blend...  ...automate and modernize workflows across eDiscovery, CLM, and AI platforms. You will own building AI-powered workflows, connecting systems... 
    Platform

    Socket

    San Francisco, CA
    5 days ago
  • $172.5k - $260.1k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...a hands-on Partner Technical Architect to help top SI...  ...quickly into the Salesforce platform.What You'll Actually...  ...agents, automations, internal tools, or technical...  ...orchestration, or automation systems.Ability to reason... 
    Platform
    Full time
    Work at office

    Salesforce

    San Francisco, CA
    3 days ago
  • $175k - $240k

     ...Notion is the collaborative AI workspace where teams...  ...building a business’s system of record to making and...  ...will lead the hands on technical delivery of our most...  ...needs and ensure Notion’s platform is ready for scale....  ...standards, scripts, and internal delivery tooling. Help... 
    Platform
    Local area

    Notion, LLC

    San Francisco, CA
    3 days ago
  •  ...ResponsibilitiesAs a Principal ML System Engineer in the Rovo & AI Engineering org, you will...  ...involved projects from technical design to launch. You will...  ...with other teams and internal customers to set expectations...  ...haveExpertise in taking a platform approach to building... 
    Platform
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  • $227.2k - $284k

    Scale's Physical AI business unit is dedicated to solving the data...  ...Physical AI.The RoleAs an ML Systems Engineer on the Physical AI team, you will design and build platforms for scalable, reliable, and...  ...production systems, supporting both internal research discovery and... 
    Platform
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed framework for large language model training and inference...  ...at the heart of the field of AI as an indispensable provider of...  ...you’d have:Strong excitement about system optimizationExperience with multi-node... 
    Platform
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...What to expect As an intern, you will be responsible for helping...  ...software and machine learning systems that power robot perception,...  ...software systems for robotic platforms that power robot perception, intelligence...  ...learning, or multimodal AI Have experience building... 
    Platform
    Internship
    Immediate start

    Human Computer Lab

    San Francisco, CA
    2 days ago
  • $224k - $308k

    Secure Every Identity, from AI to HumanIdentity is the key to...  ...Position Description:The Services Architect is a technical authority on both cloud and on-premises based IT systems and is responsible for...  ...industry leading cloud identity platform for our customers. You will... 
    Platform
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $251.84k - $314.79k

     ...transformation, to the context teams and AI systems rely on. Fivetran helps organizations...  ...About the RoleFivetran needs an engineer-architect who can make the AI analyst experience real...  ...to an innovative mental health support platform that offers personalized care and... 
    Platform
    Full time
    Work at office
    Remote work
    Flexible hours

    Fivetran

    Oakland, CA
    4 days ago
  •  ...purpose.We partner with leading platforms—Salesforce, ServiceNow,...  ...to resilient case management systems and optimized workforce planning...  ...functions, or activate AI-powered innovation, we connect...  ...Salesforce - Agentforce Technical Architect / Senior Developer Slalom is... 
    Platform
    Temporary work
    Local area

    Slalom

    San Francisco, CA
    7 hours ago
  • $148.2k - $292.3k

    Position Summary Join our AI & Engineering team in transforming technology platforms, driving innovation, and...  .../30/2026. Work you'll do As a Technical Architect on the Banking and Capital Markets...  ...designing and building systems using microservices and event-... 
    Platform
    Local area

    Deloitte

    San Francisco, CA
    1 day ago
  • $148.19k - $231.98k

     ...Salesforce Salesforce is the #1 AI CRM, where humans with...  ...6. The Role As a Core Technical Architect, you are a pre-sales,...  ...applications, delivery, and system architecture. Coding experience...  ...in business as the greatest platform for change and in companies... 
    Platform

    100 Salesforce, Inc.

    San Francisco, CA
    1 day ago
  • $154.2k - $192.8k

     ...looking for a Customer Support Systems & Analytics Architect to own the data strategy...  ...model, and sophisticated AI automation—we are moving away...  ...and projects - from our internal Support team to our BPO partners...  ...and reports for non-technical stakeholdersDeep understanding... 

    Mercury

    San Francisco, CA
    1 day ago
  • $342k

     ...Hardware organization develops system and infrastructure solutions...  ...tailored to the demands of advanced AI workloads. We work across the...  ...—partnering closely with internal teams and external vendors to...  ...About the RoleWe are seeking a 3P Architect to define and drive rack- and... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $132k - $179k

     ...making through our leading AI-infused scenario planning and analysis platform so our customers can...  ...looking for a SOLUTION ARCHITECT who delights in the most...  ...skillsDemonstrated knowledge of a formal system implementation...  ...by a member of our internal recruitment team whenever... 
    Platform

    Anaplan

    San Francisco, CA
    4 days ago
  • $210k - $256.67k

     ...ready?The TeamOur Digital Platforms Practice helps some of...  ...data migration)Ensure that technical solutions, platforms, and systems satisfy business requirements...  ..., API Management, AI, Advanced Analytics and IOT...  ...vendors as needed.Manage internal consulting resources, 3rd... 
    Platform
    Full time
    Contract work
    Temporary work
    Work experience placement

    Infosys Technologies

    San Francisco, CA
    3 days ago
  • $199k - $273.9k

     ...Every Identity, from AI to HumanIdentity...  ...secure, scalable systems that power how our...  ...and ensuring these platforms evolve alongside...  ...and we are actively architecting the next...  ...will serve as the technical anchor for Okta's...  ...and oversight for internal delivery teams and... 
    Platform
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $151k - $204.3k

     ...grow.As a Solutions Architect at AWS, you'll build deep technical relationships with...  ...latest innovations in AI and machine...  ...multi-step workflows. Internally, you will be the voice...  ...adopted cloud platform. We pioneered cloud...  ..., cloud computing, systems engineering, infrastructure... 
    Platform
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    4 days ago
  • $264.8k - $331k

    AI is becoming vitally important in every function of our society...  ...next-gen Agent RL training platform, support large scale training,...  ...technologies to optimize our ML system. Your customer will be other...  ...the art models, developed both internally and from the community, to define... 
    Platform
    Full time

    Scale AI

    Daly City, CA
    4 days ago
  • $160k - $180k

     ...: We’re a team of AI, technology, and language...  ...Generation (RAG) platform. Our proprietary,...  ....As a Solutions Architect at Pryon, you will...  ...needs into robust technical solutions. From...  ...requirements.Contribute to internal knowledge bases,...  ...enterprise-grade systems for federal or... 
    Platform
    Full time
    Temporary work
    Remote work

    Pryon

    San Francisco, CA
    4 days ago
  • $171k - $250.8k

     ...a Business Solutions Architect / Product Manager responsible...  ...for Axon’s Finance Systems portfolio. This role...  ...for critical finance platforms that support close, consolidation...  ...delivered and why; technical teams own how it is...  ...efficiency and internal controls.Support technical... 
    Platform
    Work experience placement

    Axon

    San Francisco, CA
    3 days ago
  •  ...technologies. The Global Systems Integration team has a...  ...partner with leading platforms—Salesforce, ServiceNow,...  ...functions, or activate AI-powered innovation, we...  ...platforms. * Create detailed technical specifications,...  ...contributions (e.g., accelerators, internal IP, client POVs)... 
    Platform
    Temporary work
    Local area

    Slalom

    San Francisco, CA
    4 days ago
  • $189k - $274k

     ...This role builds the systems that make it work...  ...Content Systems Architect, you’re responsible...  ...’s editorial and AI-generated language...  ...models, agents, and platform integrations that...  ...across Adobe’s internal AI platforms and tools...  .... Define the technical and system requirements... 
    Platform
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    10 days ago
  • $154k - $208k

     ...rebuilding biotech for the AI era.When a...  ...is the AI platform for biotech R&D. Scientists...  ...Solutions Architect with deep expertise...  ...enterprise software systems to partner with Benchling...  ...advisor and technical leader for our enterprise...  ...partnership with internal engineering,... 
    Platform
    Work at office
    Local area
    Flexible hours
    3 days per week

    Benchling

    San Francisco, CA
    7 hours ago
  • $151k - $204.3k

     ...state-of-the-art AI/ML/GenAI tools on...  ...including Agentic platforms, tool usage, evaluations...  ...Solutions Architect (ML SA), who will...  ...You must have deep technical experience working...  ...and support an AWS internal community of GenAI...  ...cloud computing, systems engineering, infrastructure... 
    Platform
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    3 days ago
  • $143k - $190.4k

     ...product experts. The technical solutions team...  ...Partner Solution Architect (PSA) is a Datadog...  ...teams;Collaborate internally to support Partner...  ...practices for Datadog platform adoption and...  ...channel partners (Systems Integrators, Managed...  ...platform for the AI era, providing businesses... 
    Platform
    Work at office
    Worldwide

    Datadog

    San Francisco, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems & Platform Internals - Technical Architect. Be the first to apply!