AI Infrastructure Engineer
Percepta
AI Infrastructure Engineer
Percepta's mission is to transform critical institutions with applied AI. We care that industries that power the world (e.g. healthcare, manufacturing, energy) benefit from frontier technology.
To make that happen, we embed with industry-leading customers to drive AI transformation. We bring together:
- Forward-deployed expertise in engineering, product, and research
- Mosaic, our in-house toolkit for rapidly deploying agentic workflows
- Strategic partnerships with Anthropic, McKinsey, AWS, companies within the General Catalyst portfolio, and more
Our team is a quickly growing group of Applied AI Engineers, Embedded Product Managers and Researchers motivated by diffusing the promise of AI into improvements we can feel in our day to day lives.
Percepta is a direct partnership with General Catalyst, a global transformation and investment company.
About the Role
We're hiring an AI Infrastructure Engineer to own the infrastructure, deployment, and operational reliability that powers Percepta's AI systems, including the autonomous agents at the core of what we ship.
Part of the work is hardening what exists: tightening our Terraform footprint, strengthening deployment pipelines, bringing more rigor to how we manage infrastructure across regions and providers. Part of it is building what's missing. And part of it is genuinely new territory, figuring out what SRE means when the systems you're operating make autonomous decisions.
The infrastructure patterns for the agentic systems of the future don't exist yet. You'll help define them.
Why This Is Different
- You're deploying autonomous systems. The infrastructure contract changes when your workloads have agency.
- Observability means understanding why an agent made a decision, not just whether a pod is healthy.
- The gap between research and production is real here. Our teams move optimization algorithms and AI systems from research environments into production, and you'll be part of that handoff. MLOps experience isn't required, but you'll be closer to that boundary than most infra roles.
- Small team. Real ownership. You're making foundational decisions, not inheriting someone else's.
What You'll Do
- Define infrastructure patterns for multi-agent systems that need to be observable, controllable, and recoverable in ways traditional apps don't require
- Own and evolve our IaC stack: Terraform and Kubernetes across AWS, GCP, and Azure
- Build observability primitives for agentic workflows, tracing agent decisions and execution paths, not just service latency and pod health
- Design and maintain CI/CD pipelines that give teams fast, trustworthy feedback from commit to production
- Build operational foundations: monitoring, alerting, incident response, and the new patterns that emerge when AI systems are participants in that response
- Work across engineering teams to meet the reliability and compliance requirements of the institutions we serve (SOC 2, HIPAA, regulated environments in healthcare and energy)
What We're Looking For
- 5+ years building and operating production infrastructure in DevOps or SRE roles
- The kind of engineer who sees a manual process and can't rest until it's automated well, not just scripted
- Strong hands-on Terraform experience
- Deep experience with at least 1 major cloud provider (AWS, GCP, or Azure): networking, IAM, cost management, the operational realities of production workloads
- Solid Docker and Kubernetes experience in production. We run managed clusters across all 3 major clouds; this is a core part of the role
- Experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, or similar)
- Scripting proficiency in Python, Bash, or similar
- High agency: you don't wait for a ticket to fix what's broken, but you communicate, collaborate, and bring the team along
- Genuine curiosity about AI systems, not just the infrastructure running them. You want to understand what you're operating
- You find it interesting (not alarming) that some systems you'll operate will be making decisions on their own
Nice To Have
- Multi-region and multi-cloud experience across 2+ providers
- Experience with single-tenant or on-prem deployments alongside multi-tenant SaaS
- Familiarity with GitOps patterns and progressive delivery
- Familiarity with the Grafana stack (Prometheus, Grafana, Loki) or equivalent
- Experience with compliance frameworks (HIPAA, SOC 2) and how they shape infrastructure decisions in regulated environments
- Background supporting ML or research workflows moving to production: model deployment, pipeline orchestration, or similar
- You've thought about what observability means for non-deterministic systems and have opinions about it
The infrastructure patterns for autonomous AI systems are still being written. If you want to be one of the people writing them, let's talk.
We're working against an incredibly ambitious mission. It won't be easy, but it will likely be the most fulfilling work of your career. If this excites you, let's chat, even if you don't meet all of the qualifications above.
Our Values
Dream bigger: We have the unique privilege of taking on the most ambitious problems and we should chase them with optimism, responsibility, and genuine belief that we can make it happen. We have to embrace the hard things when no one else will. Heart in the game: What we're doing matters and we have to give a shit. Internally, that means fixing badness when you find it. Externally, it means honoring the trust our customers place in us with their most important problems. This isn't a 9-5, nor is it a job we're ever going to monitor your hours. We promise to put work in front of you that matters and in return, we ask you to promise to care. Win for the customer: Everyone is an engineer and the job of an engineer is to deliver outcomes, not outputs. Everything we do—the products we build, the partnerships we launch, the strategy we set—exists to make our customers successful. Delivery is the strategy. Make the call: Organizations are only as strong as the pace at which they make decisions. Everyone at Percepta should feel empowered to commit and shape the ambiguity in front of them. But "make the call" cuts both ways: make the decision and make the phone call. High-agency decision-making only works with high-bandwidth communication and we commit to never operate in silos. Intensity with kindness: We believe in excellence in execution, candor in feedback, ruthlessness in prioritization, and survivalist urgency. We also believe you don't need to be an asshole to deliver on any of this. The trust built through shared kindness and vulnerability is what makes the intensity sustainable.
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered... ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...SuggestedFull timeWork experience placementLive inWork at officeLocal area
$215k - $350k
We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering...SuggestedWorldwideHome office$175k - $275k
...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in... ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,...SuggestedFull timeWork at officeFlexible hours- ...Role Overview: Primary hiring focus is an AI Infrastructure Senior Engineer supporting the build-out of the company's Azure-based technology stack. The role will be the first hire on the infrastructure engineering team and will work under the infrastructure...SuggestedCurrently hiringWork at officeRemote work
$200k
...AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role...SuggestedWork at office$150k - $300k
...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in... ...constantly. This is a place to grow your career, make a real impact, and help define a new category of infrastructure software....Full timeWork at officeFlexible hours- ...About Build AI for the Built World: Build has created the agentic AI stack for institutional... ...most important built projects - digital infrastructure, energy, industrial - from concept to... ...create and configure workflows without engineering. The interfaces through which clients...Full timeLive in
$160k - $220k
...of: The role We are looking for an experienced AI Engineer to lead the implementation of Azure AI Foundry within an established... ...our existing data platform, governance model, and analytics infrastructure. You will work closely with data engineering,...Full timeRemote work$229.9k - $262.4k
Senior Lead AI Engineer, Gen AI Platform Overview: At Capital One, we are creating responsible and reliable... ...personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in...Full timePart timeLocal area- ...Description Job Description Palona’s AI agents operate continuously in... ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...Temporary work
$162k - $215k
AI-Enabled Infrastructure And Systems Engineer Aladdin Platform Engineering powers the technology foundation behind BlackRock's Aladdin platform - a unified system that connects risk, portfolio management, trading, and operations for investors globally. We build the mission...ApprenticeshipWork at officeWork from homeWorldwideFlexible hours1 day per week$100k - $160k
AI Infrastructure Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- Senior AI Storage Infrastructure Engineer Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to...Local area
$100k - $150k
AI Infrastructure Engineer- Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- Job Title: AI Infrastructure Engineer Job Summary We are seeking an AI Infrastructure Engineer to design, implement, and manage the infrastructure that powers AI and machine learning workloads. In this role, you will build scalable, secure, and high-performance environments...Full timeRemote work
- ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and... ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations...Full timeWork experience placementLive inWork at officeLocal area
$175k - $200k
...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant... ...into production, and our external agent infrastructure (the Agency Connect MCP server, which... ...early partner adoption. Both need a senior engineering owner, and both need to grow into the...Full timeInternship- ...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco... ...are seeking a Director of AI Platform Engineering to lead the design, development, and... ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent...Ongoing contractFull timeCasual workWork at officeFlexible hours
$152.29k - $250.2k
As the Head of AI Platform Engineering - Execution Plane, you will lead the development and implementation of our enterprise platform’s execution... ...design reviews, architecture standards, testing, CI/CD, infrastructure automation, incident response, SLOs, runbooks, and...Full timeLocal areaVisa sponsorshipWork visaFlexible hours- ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits...
- ...Job Title: AI Platform Engineer Location: NYC, NY (Hybrid Model) - 10003 Zip code Energy & Utility domain with experience in Google... ...Platform Engineer: Design and implement AI/ML infrastructure using Vertex AI, Kubeflow, TensorFlow Extended (TFX), and...Local area
- ...Senior AI Platform Engineer Location: New York City, NY/Hybrid Duration: 12+ Months Experience required: 6 years Interview Mode: Video... ...cost guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject matter expert (SME) on Gen AI...
$155k - $215k
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm's Technology principles... ...of AI across our businesses.This role is for a platform engineering specialist who will help build a firmwide AI Development Platform...Temporary work$230k - $290k
...York( Digital Solutions Group ) - Platform Engineering /Full Time /RemoteAHEAD helps large... ...same enterprises design, build, and run AI agent platforms on top of it. This role... ...delivery leadership attached. You will write infrastructure code, agent code, and the documents that...Full timeWork at officeShift work$180k - $230k
About the RoleThe Platform Infrastructure team at iCapital plays a critical role in ensuring that... ...and intellectually curious MLOps/DevOps Engineers with deep expertise in machine learning... ...).Enable production workloads for AI/ML and Generative AI systems, including...Full timeWork at officeRemote work- The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data...
$205k - $235k
...With deep functional and sector expertise, paired with innovative AI-powered technology and an investor mindset, we partner with... ...Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data...Full timeFor contractorsWork experience placementSummer holidayFlexible hours- As the Lead AI Engineer for our next-generation TCO Agent Platform, you will serve as the technical lead and multi-agent system architect... ..., and API/MCP error specifications (RFC 7807/9457) DevOps & Infrastructure: Proficiency with Docker builds, GKE deployment patterns,...
- Google Cloud AI Engineer POC We are seeking a highly skilled Artificial Intelligence Engineer. This role is pivotal in establishing... ...preprocessing techniques (scaling, encoding, imputation). • Cloud Infrastructure: Hands-on experience with Google Cloud Storage and Vertex AI...
$152k - $241.5k
...weight models are foundational to American AI leadership and cybersecurity, and that... ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling... ...and maintain the agent harness.Evaluation infrastructure: Build the systems we use to run and...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
- ai engineer New York, NY
- ai research engineer New York, NY
- ai engineer remote New York, NY
- ai prompt engineer New York, NY
- ai developer New York, NY
- senior ai engineer New York, NY
- machine learning ai engineer New York, NY
- ai ml engineer New York, NY
- lead infrastructure engineer New York, NY
- principal infrastructure engineer New York, NY



