AI Systems & Platform Internals - Technical Architect
Accellor
Accelloris an AI-native services firm purpose-built for the post-ChatGPT era. Free from legacy constraints, we focus on delivering measurable business outcomes through advanced AI, data, and engineering capabilities.Ourmission isto operationalize AI at scale and unlock sustained enterprise value. Our offerings spanAI solutions, data services, enterprise applications, and product engineering, tailored to industry-specific needs across healthcare, life sciences, telecom, retail, financial services, and technology. Byleveragingdesignthinking and technology-agnostic architectures, we ensure faster time-to-value and seamless interoperability. With a proventrack recordof enabling Fortune 100 enterprises and global innovators,Accellorstands as a trusted partner for organizations seeking to harness the full potential of AI. Our vision is clear:to build intelligent, connected ecosystems that deliver measurable outcomes and redefine the future of enterprise transformation. Technical Architect — AI Systems & Platform Internals Experience: 10–12 Years Role Type: Technical Architect / Staff-Level Systems Architect Role Summary Accellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power ChatGPT, OpenAI API, Codex, agentic systems, multimodal experiences, and internal research workloads. This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability. The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations. Key Responsibilities 1. AI Systems Architecture Design and evolve large-scale AI systems that support ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads. Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout. Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency. 2. Inference Runtime & Model Serving Architect high-throughput, low-latency inference systems across large-scale GPU clusters. Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization. Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request. Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving. 3. GPU, Kernel & Distributed Performance Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling. Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution. Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement. 4. Context Engineering Design and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured. Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability. Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality. 5. Cost Optimization Frameworks Design and build cost optimization frameworks for large-scale LLM and GenAI workloads. Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend. Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows. Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience. 6. Training & Research Infrastructure Collaborate with research and training infrastructure teams to support large-scale model training and post-training workflows. Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity. Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution. 7. Release Safety, Validation & Evaluation Gates Architect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free. Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality. Ensure platform optimizations do not reduce safety, grounding, quality, or user trust. 8. Reliability, Observability & Production Operations Design systems that make AI infrastructure observable, debuggable, reliable, and operationally safe. Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops. Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures. Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure. 9. Agentic & Multimodal Platform Internals Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment. Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints. Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads. 10. Technical Leadership Work closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams. Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions. Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs. Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries. Required Qualifications 10–12 years of experience in software engineering, systems architecture, ML infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering. Strong hands-on engineering experience with Python and at least one systems/backend language such as C++, Go, Rust, Java, or TypeScript . Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture. Experience designing or operating large-scale systems involving APIs, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring. Strong understanding of AI/ML systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment. Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL. Experience with ML frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems. Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services. Strong communication skills with the ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions. Preferred Qualifications Experience working on LLM inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms. Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving. Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or LLM cost dashboards. Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems. Experience with large-scale distributed training, RL infrastructure, checkpointing, ML compiler optimizations, model graph transformations, or training runtime systems. Experience designing release gates, regression detection systems, canary systems, CI/CD validation frameworks, and production safety controls for performance-sensitive infrastructure. Experience with evals, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring. Technical Skill Areas AI Systems: LLM serving, inference runtime, training infrastructure, post-training workflows, agent systems, multimodal models Inference: batching, routing, KV-cache, streaming, latency optimization, model serving, tensor parallelism, pipeline parallelism Performance Engineering: CUDA, Triton, GPU profiling, kernel optimization, memory bandwidth, communication libraries, distributed execution Context Engineering: prompt architecture, dynamic context assembly, RAG, memory, context compression, context ranking, source grounding, permission-aware retrieval Cost Optimization: token budgeting, caching, model routing, fallback strategies, cost telemetry, batching, async workflows, cost-quality trade-offs Distributed Systems: scheduling, orchestration, reliability, fault tolerance, observability, scalability, service design ML Frameworks: PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Ray Infrastructure: Kubernetes, Docker, Terraform, CI/CD, cloud platforms, Linux systems, networking, storage Safety & Validation: evals, release gates, canaries, regression testing, model behavior validation, rollout safety Candidate Profile The ideal candidate is a senior hands-on architect who can operate across the full AI systems stack. They should be able to discuss GPU memory bottlenecks, distributed inference, model-serving reliability, context quality, cost optimization, release validation, eval pipelines, observability, and production rollout with engineering teams, while also explaining architecture decisions clearly to senior leadership. The candidate should not be limited to architecture diagrams. They must be capable of reviewing implementation quality, identifying bottlenecks, debugging production issues, challenging weak assumptions, and converting repeated failures into stronger platform abstractions. This role requires the judgment of a senior architect, the debugging mindset of a systems engineer, and the ownership mindset required for production AI infrastructure. #J-18808-Ljbffr Accellor
- ...an alert: Customer Systems Architect - Data Center Solutions... ...factory-built product platforms. We are seeking a high-gravity, technically authoritative Customer... ...hyperscale clients and our internal product development... ...center infrastructure for AI, cloud, and hybrid...PlatformWork at officeLocal areaRemote workWorldwideShift work
$184k - $287.5k
System Architect page is loaded## System Architectlocations:... ...state of the art compute platforms for the world to use.... ....* Demonstrate technical leadership in power architecture... ....* Align with internal stakeholders on technical... ...existing vacancy.NVIDIA uses AI tools in its...PlatformTemporary work- ...Cloud, is a leader in AI cloud... ...translating emerging platform capabilities into... ...rack and pod level system arrangement that maximizes... ...HPC Systems Architect with extensive experience... ...Collaborate with internal teams and... ...implementation teams. Provide technical leadership and...PlatformWork at officeLocal areaWork from homeFlexible hours
- ...Development, the Principal Systems Architect - Connected Devices... ...for ensuring the platform is scalable, secure, and... ...in collaboration with internal teams and external development... ...partners. As the technical authority and system... ...teams. Exposure to AI-assisted development tools...PlatformFull timeRemote workWorldwide
$250k - $365k
...interpretable, and steerable AI systems. We want AI to be safe and... ...them how. That's you. Technical Architects help customers get real work... ...webinar one day, a customer's platform team in their own... ...required will correlate with the internal job level requirements for...PlatformWork at officeVisa sponsorshipFlexible hoursShift workDay shift$141.4k - $215.4k
...software. As a Solution Builder Lead System Architect, you will serve as the senior technical advisor on client engagements,... ...operations through the Pega Platform. You are more than a technical architect... ..., Pega Infinity, Pega Cloud, and AI-assisted development....PlatformLocal areaRemote workFlexible hours$113.8k - $176.5k
...member of the Pega APAC System Architecture team, you... ...innovative business and technical solutions using Pega... ...(PDC) Mentor System Architects and act as a trusted advisor... ...problems, with deep platform expertise and specialist... ...input continues] ... AI in Action – Responsible...PlatformRemote workFlexible hours- Neurophos, Inc. in Austin, TX, seeks a Systems Engineer to own the OVMM engine architecture and translate product-level requirements into... ...combines architecture, requirements, and leadership to deliver a first-of-a-kind computing platform. #J-18808-Ljbffr Neurophos, Inc.Platform
- ...DIVISION IS SEEKING A SYNTHETIC SYSTEMS ARCHITECT TO LEAD THE DESIGN AND... ...ARCHITECTURES FOR OUR SYNTHETIC PERSON PLATFORM. THIS ROLE SITS AT THE... ...NEUROSCIENCE, ADVANCED AI, AND SYSTEMS ENGINEERING. RESPONSIBILITIES... ...EXPERIENCE LEADING TECHNICAL TEAMS OF 10+ ENGINEERS...Platform
$139.2k - $187.05k
...to scale to gigawatts for next-generation AI data centers. Boom Supersonic is... ...keep up. This role owns the medium-voltage system that turns shaft power into dependable 13... ...engineering at industrial scale. The Superpower platform is designed to operate in arrays that...PlatformPermanent employmentFlexible hours$204k - $216k
Sapience AI is the collective intelligence platform for professional communities. We sit above... ...in silos, in legacy systems, in the heads of a few experts... ...of the system, the big technical bets, and the seams... ...The Principal AI Systems Architect owns that. You set cross...PlatformFor contractors- ...scientific discovery to powering AI and the technologies people... ...dynamic, energetic Lead / Principal Systems Design Engineer to join our... ...and encourages continuous technical innovation to showcase successes... ...to product architecture, platform definition, chipset, firmware...Platform
$173.46k - $231.98k
...Salesforce is the #1 AI CRM, where humans with... .... The Data & AI Cloud Technical Architect The Data & AI Cloud Technical... ...combines deep data platform knowledge with broad... ...to agentic AI systems and LLM‑powered applications... ...customers, partners, and internal audiences. Baseline...Platform- ...Pencils is seeking a seasoned AWS AI Solutions Architect to lead the design and... ...grade generative and agentic AI systems on AWS. You will architect scalable, secure platforms using Bedrock, Bedrock... ...related services. As a strategic technical advisor, you translate ambiguity...Platform
- Lambda is seeking a Senior Business Systems Architect to lead enterprise systems architecture across financials, supply chain, and customer platforms. You will design end-to-end solutions, manage complex integrations, and drive a strategic roadmap for system enhancements...Platform
$140k - $180k
## Solutions Architect - Digital Systems-To all recruitment agencies: Formlabs does... ..., you will be a strategic technical leader within the Systems Department... ...focusing on how our core platforms (Salesforce, NetSuite,... ...of new toolsets, including AI supported workflows and...PlatformFull timeWork at officeFlexible hours3 days per week$115k - $200k
...distributors around the world. AI at SharkNinja At... ...of manual work, or a system that doesn’t yet exist... ...success must look like. Architect and build production-grade... ..., algorithms, data platforms, integrations, and AI agents... .... Who You Are Deeply technical software and systems...PlatformTemporary workLocal areaFlexible hours- ...numerical Python, not an internal tool that happens to be... ...to keep the whole system coherent as it grows.... ...off users and into the platform. Primitives. Cholesky,... ...can vouch for expand, technical debt goes down, the core... ...it in and ship it. Use AI agents if they help. They...Platform
- ...Solutions Architect United States | Hybrid Supervity AI is building a new enterprise... ...to be the chief technical authority for our... ...architectures, multi-agent systems, and the... ...customers using our platform. This includes designing... ...customers and internal teams. You will...PlatformLocal areaRemote work
- ...category leader in AI-native readiness... ...impact, inspired by technical depth, and ready... ...looking for a Solutions Architect to join our Go-To-... ...stakeholders and internal teams,translating... ...national security systems and environments.... ...solutions, cloud platforms, and/or systems...PlatformRemote workFlexible hoursShift work
$165k - $185k
...Senior Solutions Architect to play a critical... ...opportunities where technical feasibility,... ...both customers and internal teams navigate sophisticated... ...with enterprise systems such as Workday,... ...to leverage AI to better support... ...driven team building a platform that solves real...Platform$107.5k - $143.3k
...Description: Solution Architect Role Overview Are... ...cloud and AI? As a Solution Architect... ...actively test platform limits, build reusable... ...: Guarantee system throughput and p95... ...practices. • Pre-Sales Technical Execution: Partner... ...learning resources, and internal advancement...PlatformMinimum wageRemote workFlexible hours- ...experienced Anaplan Solution Architect to work closely with... ...advanced planning systems like Anaplan. In this role... ...Use leading planning platforms to empower decision-makers... ...to bottom-up internal initiatives or bringing... ...advanced optimization and AI forecasting to fit client...PlatformWork experience placementFlexible hoursWeekend work
- ...:****IT Solutions Architect – WDM Iowa**We are... ...architecture and technical direction for key... ...technologies (InsurTech), AI-enabled solutions,... ...tools, and cloud platforms, and recommend... ...with clients (internal and external) by supporting... ...patterns, system design principles,...PlatformFull timeTemporary work
$180.2k - $247.7k
...combining frontier agentic AI, an enterprise-grade platform, and deep domain... ...Senior Solutions Architect at is a critical role... ...work directly with technical and non-technical stakeholders... ...customers’ existing systems, workflows, and data... ...product roadmap and internal priorities based on...Platform$160k - $190k
...DuploCloud is an Agentic Internal Developer Platform (IDP) for DevOps and... ...use case as an AI agent and deploy it to... ...looking for a Solutions Architect with a demonstrated... ...is a customer-facing technical pre-sales role: you will... ..., and of CI/CD systems (e.g., GitHub Actions...PlatformPermanent employmentWork experience placementLive inWork at officeLocal areaRemote workFlexible hours$112.9k - $257k
...AI Solution Architect The Opportunity: Design and implement... ...in both company and technical competencies. You... ...mission or enterprise systems, including data... ...application workflows on platforms such as Palantir... ...federal, state, local, or international law. Know Your...PlatformFull timeContract workPart timeWork at officeLocal areaRemote workShift work$100k
...software that helps architects, engineers, and manufacturers... ...driven and technical expert professional with... ...presales meetings both internally and with customers,... ...specification of appropriate system platforms, integrations and... .... Graitec uses AI to support and...PlatformWorldwideFlexible hours$150k - $180k
...HealthEdge® offers AI-powered operational infrastructure... ....com. The Solution Architect will help in... ...third party vendors, and internal teams to define and... ...solutions and support technical system implementations. Candidate... ...Core admin platforms (Claims adjudication,...PlatformPermanent employmentFull timeWork experience placementWork at officeRemote work$152k - $192k
...*Senior Solutions Architect****Hybrid role (3... ...sustainable health care system.****Who We Are... ...business and technical problems. They lead... ...and future-state platform design across critical... ...solutions that leverage AI, automation,... ...skills with both internal and external groups...PlatformWork at officeImmediate startWork from homeRelocationFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Systems & Platform Internals - Technical Architect. Be the first to apply!
- technical architect Eastern, KY
- servicenow technical architect Eastern, KY
- pega system architect Eastern, KY
- embedded systems architect Eastern, KY
- system architect Eastern, KY
- salesforce technical architect Eastern, KY
- power platform Eastern, KY
- platform manager Eastern, KY
- platform product manager Eastern, KY
- director of digital platform Eastern, KY

