Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Platform & Agentic Infrastructure Engineer

$178k - $321k

OKX

Who We Are At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me, Do the Right Thing, and Get Things Done. These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more.About the OpportunityOKX’s Internal Audit function has an early but working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data pipelines, and AI-enabled tools that let a very small team punch far above its weight. Your job is not to rebuild it at today’s maturity. Your job is to take it to a regulator-grade production standard and well beyond, and to drive AI Enablement across the department. You own the foundation, the agent runtime, and the harness: a resilient cloud platform, the agentic runtime and evaluation harnesses that make agents trustworthy in a regulated setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated building, not on claims or credentials. We expect you to be more capable than the hiring manager in your domain: you will co-own and challenge the tooling strategy, not just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your counterpart’s stack when needed. EnvironmentCloud on AWS or GCP. Claude as the starting point in a deliberately multi-model architecture (Anthropic, OpenAI, Google), using the right model for each job, including the plugins and integrations you build. Google Workspace for reporting and evidence. JIRA for the audit team’s workflow. Lark for team communications, alert bots, and the corporate wiki. You inherit a working Google-native prototype estate (Apps Script web apps, Drive-synced automation, locally scheduled jobs, Claude Code agent tooling) and evolve it without breaking daily use. Infrastructure-as-code, CI/CD, and observability throughout. Treat this stack as the starting point, not a constraint: you build production-grade systems end to end on what exists today, and you are expected to propose, prove, and adopt better components as demand and capabilities evolve. Production-grade here means service levels sized for an internal assurance platform: board-cycle windows are sacred, recovery is measured in hours, and this is not a 24/7 pager culture.In your first yearHive Mind runs in the cloud with HA, DR, SLOs, and audit logging that passes an internal controls review.A governed data foundation with provenance and lineage is live across multiple audit domains, integrated with OKX's group data infrastructure where it exists.An agentic runtime and harness with evaluation and red-teaming gates what reaches production.The hiring manager is out of the operational loop: no scheduled job runs on a personal machine, every system has a runbook and a non-founder owner. Decommissioning the founder’s laptop as infrastructure is a literal milestone. When these goals compete, the priority order is: keep the estate alive, then the cloud migration with observability, then the harness gating production, then the data foundation.How we assess demonstrated buildingWe assess demonstrated building in ways that respect your time and your confidentiality obligations to current and former employers: a portfolio deep-dive (walk us through systems you built and kept running, at the level of detail your obligations allow; we want architecture, decisions, and trade-offs, never proprietary code, data, or documents), a short, time-capped build exercise on a synthetic problem unrelated to OKX's business (a small agentic workflow with an evaluation harness, used for assessment only and never put to use by OKX; the work remains yours) that you defend live, walking us through your design decisions and extending it on the spot, a systems-design session on taking a prototype estate to production, and references focused on whether you built and operated systems in production.Trust and complianceThis role handles highly sensitive audit data at a global crypto exchange. Expect background checks, confidentiality obligations, and personal-trading and material-non-public-informationWhat You’ll Be DoingInherit, operate, and progressively migrate the working prototype estate (Google Workspace–native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows. Working software wins arguments; migrate by strangling, not rewriting. Reuse and integrate with OKX's prevailing and evolving AI capabilities and data infrastructure, including enterprise-approved models and gateways, Model Context Protocol (MCP) servers and connectors, security tooling, and group data platforms, before building parallel capability.Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production.Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content.Build the Responsible-AI and model-governance layer: hallucination, bias, and drift controls, output validation, guardrails, and complete logging. Extend the existing governance design (human-calibrated judge gates, golden-set regression, provenance registry) rather than replacing it. New model providers enter through the same governance and approved-tooling review, not around it.Build data infrastructure with provenance and lineage: immutable audit trails, versioned evidence, reproducible pipelines, and traceability from source to report.Engineer data protection: encryption, key and secrets management, least-privilege access, sensitive-data handling, residency, and defensible retention.Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation: analytics over full populations, with exceptions streamed in real time. Auditors and the hiring manager define the audit logic; you make it run at production grade.Own the cloud foundation: infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, and make the platform examinable.What We Look For In You 7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time.Proven brownfield migrations: you have taken a founder-built or prototype system to production grade while it stayed in daily use.Strong engineering fundamentals: data structures and algorithms, fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go), with the full-stack range to connect the infrastructure yourself.Agentic runtime and harness engineering, proven by building: the field is too young to demand years of it, so we weigh real systems shipped over tenure. You have genuinely built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill- and hook-based agent tooling, and the evaluation and red-team harnesses that grade agent behavior, with responsible-AI controls (hallucination, bias, drift).Data engineering with provenance and lineage, and security and data-protection engineering by default (encryption, identity and access management, secrets, retention). You build systems that could withstand external audit, by design.Cloud (AWS or GCP) plus resilience engineering: infrastructure-as-code (Terraform), CI/CD, HA, DR, SLOs, and observability, backed by automated testing and documentation.You ship inside locked-down enterprise environments (TLS-intercepting proxies, endpoint detection and response tooling, restricted installs, security guardrails, OAuth admin consent) without treating security as someone else’s problem.Partnership: you translate audit needs into systems, explain technical risk to non-engineers, and drive adoption.Nice to HavesMulti-agent orchestration frameworks, MCP servers, and tooling and plugin development across the Anthropic, OpenAI, and Google model ecosystems.Model-risk or AI-governance program experience; frameworks such as the NIST AI Risk Management Framework.Continuous-auditing or continuous-controls-monitoring platforms; streaming and real-time data at scale; statistical anomaly detection and applied machine learning beyond LLMs; vector stores and retrieval; LLM cost engineering.Experience passing external audit, SOC 2, or SOX (helpful, not required).Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience.Perks & Benefits Competitive total compensation packageL&D programs and Education subsidy for employees' growth and developmentVarious team building programs and company eventsWellness and meal allowancesComprehensive healthcare schemes for employees and dependantsMore that we love to tell you along the process!OKX Statement:OKX is committed to equal employment opportunities regardless of race, color, genetic information, creed, religion, sex, sexual orientation, gender identity, lawful alien status, national origin, age, marital status, and non-job related physical or mental disability, or protected veteran status. Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.The salary range for this position is $178,000 - $321,000The salary offered depends on a variety of factors, including job-related knowledge, skills, experience, and market location. In addition to the salary, a performance bonus and long-term incentives may be provided as part of the compensation package, as well as a full range of medical, financial, and/or other benefits, dependent on the position offered. Applicants should apply via OKX internal or external careers site.Notice:All official OKX vacancies are published on this website.While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website.Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX's Candidate Privacy Notice.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior AI Platform & Agentic Infrastructure Engineer in San Jose, CA vacancy
  • $174.72k - $295.68k

     ...integrating advanced AI and autonomous driving...  ....You will be a senior engineer on the team building...  ...internal AI engineering platform — the systems, services...  ...operational systems. The agentic layer (NL intake,...  ...platform services and infrastructure with strong reliability... 
    Platform
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience... 
    Platform
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $152k - $241.5k

    Join NVIDIA's Isaac Applications Engineering team and help build the platform for Physical AI robots — assembling the full...  ...it, test for it, and build the infrastructure that keeps it true as the stack...  ...integrate skill evaluations so agentic workflows are measured, not just... 
    Platform
    Senior
    Full time
    Live in
    Night shift

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next era...  ...of NVIDIA's software engineers worldwide. We collaborate...  ...challenges in infrastructure such as Kubernetes, job...  ...automated recovery.Create agentic workflows for infrastructureCollaborate...  ...for heterogeneous platforms.Experience working in... 
    Platform
    Senior
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to new markets! This role...  ...middleware, and application layers.Build and evolve agentic AI tools and frameworks to accelerate bring-up,... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are looking for a senior systems software engineer to improve the operation and...  ...of distributed system infrastructure using AI. We are passionate about...  ...of the NVIDIA software platform! We want to develop next...  ...the crowd:Experience with Agentic AI frameworks and patterns... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $264.5k

     ...seeking outstanding Senior Design Verification Engineers with a specialty in...  ...supercomputers.Our DV infrastructure and methodology team...  ...applying AI to solve problems or implementing agentic AI flowsNVIDIA is widely...  ...effective computing platform driving our success... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148k - $224.25k

     ...unlimited potential of AI to define the next era...  ...operate inference and agentic workloads. This is a new...  ...Manager to drive NVCF platform features and developer...  ...roadmap.Collaborate with engineering on feature design, prioritization...  ...to connect with senior technical customers and... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is...  ...GPUs to power a revolution in AI, enabling breakthroughs in...  ...Generative AI inference platform to make design and deployment...  ...these capabilities to support agentic inference workloads,... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $200k - $322k

     ...learning ignited modern AI — the next era of...  ....Design-for-X Engineering at NVIDIA works...  ...ll be doing:As a senior member in our...  ...as part of the AI Infrastructure requirements at an...  ...agents and multi-agentic ecosystemsExperience...  ...with cloud platforms (AWS, Azure, GCP)... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k - $322k

     ...the unlimited potential of AI to define the next era of...  ...impact on the world.This Senior Staff Client Platform Engineer role is a high-impact...  ...build scalable, automated, agentic ecosystems that define the...  ...of global endpoint infrastructure. In this role, you will set... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

     ...tapping into the unlimited potential of AI to define the next era of computing...  ...on the world.NVIDIA is hiring a Senior Staff Engineer, Enterprise SaaS Platform & Automation — a role at the...  ...OpenAI Codex, Claude, and emerging agentic systems before they reach the broader... 
    Platform
    Senior
    Full time
    Contract work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $200k - $322k

     ...of the most advanced AI builders in the world...  ...performance, and long-term platform success. The work is...  ...translating complex infrastructure into practical outcomes...  ....gainsightWork across Engineering, Product, Operations,...  ...automations, internal tools, or agentic workflows that improve... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...powers innovative AI research and developers...  ...building the AI/ML platform for improving...  ...developing scalable AI infrastructure services globally....  ...infrastructure software engineer to join our team....  ..., fine-tuning, and Agentic AI in production.As a senior DGX Cloud AI... 
    Platform
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $262k - $365k

     ...a talented team of software engineers, fostering an open environment...  ...delivery of large-scale, infrastructure systems within a changing environment...  ...across all of Google Cloud Platform (GCP).In this role, you will...  ...services in an increasingly agentic world.Google Cloud... 
    Platform
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ...across hybrid and multi-cloud infrastructure. We are building the next...  ...the hardest problems in AI: storage, access,...  ...managers, internal AI teams, platform teams, and partner engineering teams to understand requirements...  ...AI-assisted and agentic development workflows, while... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $300k - $425k

     ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...we’re actively looking for a Senior Software Engineer, Content Platform who can drive...  ...time off.How will I use AI at Roku?At Roku, we don’t just...  ...have built fluency across the agentic engineering toolchain —... 
    Platform
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  • $328.5k

     ...TVRoku is the #1 TV streaming platform in the U.S., Canada, and...  ...highly technical, strategic Senior Engineering Manager to lead our Core Content...  ...machine learning and AI capabilities. This is a high...  ...have built fluency across the agentic engineering toolchain — coding... 
    Platform
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next era...  ...are now looking for a senior software engineer to join our Hardware Infrastructure team! Our team is responsible...  ...highly available platform services and performant...  ...guidelines around adopting agentic development to improve... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $151.8k - $265.35k

     ...exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity across...  ...intelligent, adaptive, and agentic experiences to Adobe Express....  ...adaptive AI systems.Mentor senior engineers in modern AI system... 
    Platform
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    4 days ago
  •  ...building the foundation for physical AI — a unified platform that combines high-quality robotic...  ...The Role We are looking for a Senior AI Engineer to design, build, and ship AI-...  ...across the full stack — from the agentic infrastructure that powers our robot operations,... 
    Platform
    Senior
    Full time

    Dexmate

    Santa Clara, CA
    15 hours ago
  •  ...Kai is the AI company rebuilding cybersecurity for the machine...  ...bottlenecks. The Kai Agentic Platform replaces fragmented, human-...  ...leadership team: Our Heads of AI, Engineering, and Product bring...  ...applied AI scientists and an AI infrastructure engineer to transform code... 
    Platform
    Senior

    Kai

    San Jose, CA
    3 days ago
  • $196k - $310.5k

    NVIDIA is the leader in AI, machine learning, and datacenter acceleration. NVIDIA...  ...for a highly motivated and dedicated Senior DFT Infrastructure Engineer to join our DFX group. You will join...  ...and groundbreaking compute platforms for global use. Our work helps scientists... 
    Platform
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    11 hours ago
  • Define and implement the Agentic Commerce strategy, aligning with company...  ...and capitalize on emerging AI trends in automated shopping,...  ...transactions across digital platforms. Required 6+ years of experience...  ...AI/ML models, personalization engines, and automation tools. Strong... 
    Platform
    Senior
    Full time

    Paypal

    San Jose, CA
    15 hours ago
  • $272k - $431.25k

     ...NVIDIA IT’s Enterprise AI & Automation team to...  ...expand enterprise-grade agentic AI systems at one of...  ...NVIDIA’s Enterprise AI Platform drives production AI agents...  ...results across engineering, IT, supply chain, finance...  ...candidate must grasp infrastructure aspects from... 
    Platform
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ..., GPU deep learning ignited modern AI — the next era of computing. NVIDIA...  ...intelligence.NVIDIA is seeking top-tier Senior Deep Learning Infrastructure Engineers. In this role, you will play a...  ...for a highly scalable deep learning platform.Design, build, and maintain high... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next...  ...are looking for a senior systems software engineer to advance the Continuous...  ...Delivery (CI/CD) infrastructure that powers many...  ...software on the NVIDIA platform!What you will be doing...  ...:Expertise with Agentic AI frameworks, patterns... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than...  ...intelligence via additional agentic computation.Cerebras works with...  ...a highly skilled WAN Network Engineer to design, implement, manage,...  ...Cerebras:Build a breakthrough AI platform beyond the constraints of the... 
    Platform
    Senior

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer, Agentic AI! NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing...  ...implementation of sophisticated coding agents and agentic platforms capable of multi-step reasoning, self-correction, and... 
    Platform
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...and deployment of advanced AI agents and agentic systems. Architect and...  ...robust, scalable, and reliable infrastructure to support the deployment...  ..., UX designers, and other engineers to define requirements and...  .... Experience with cloud platforms (AWS) and containerization... 
    Platform
    Senior
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    15 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Platform & Agentic Infrastructure Engineer. Be the first to apply!