AI Infrastructure Engineer
$180k - $400kPercepta
Who We Are
Percepta’s mission is to transform critical institutions with applied AI. We care that industries that power the world (e.g. healthcare, manufacturing, energy) benefit from frontier technology.
To make that happen, we embed with industry-leading customers to drive AI transformation. We bring together:
Forward-deployed expertise in engineering, product, and research
Mosaic, our in-house toolkit for rapidly deploying agentic workflows
Strategic partnerships with Anthropic, McKinsey, AWS, companies within the General Catalyst portfolio, and more
Our team is a quickly growing group of Applied AI Engineers, Embedded Product Managers and Researchers motivated by diffusing the promise of AI into improvements we can feel in our day to day lives.
Percepta is a direct partnership with General Catalyst, a global transformation and investment company.
About the role
We're hiring an AI Infrastructure Engineer to own the infrastructure, deployment, and operational reliability that powers Percepta's AI systems, including the autonomous agents at the core of what we ship.
Part of the work is hardening what exists: tightening our Terraform footprint, strengthening deployment pipelines, bringing more rigor to how we manage infrastructure across regions and providers. Part of it is building what's missing. And part of it is genuinely new territory, figuring out what SRE means when the systems you're operating make autonomous decisions.
The infrastructure patterns for the agentic systems of the future don't exist yet. You'll help define them.
Why this is different
You're deploying autonomous systems. The infrastructure contract changes when your workloads have agency.
Observability means understanding why an agent made a decision , not just whether a pod is healthy.
The gap between research and production is real here. Our teams move optimization algorithms and AI systems from research environments into production, and you'll be part of that handoff. MLOps experience isn't required, but you'll be closer to that boundary than most infra roles.
Small team. Real ownership. You're making foundational decisions, not inheriting someone else's.
What you'll do
Build the operational floor: SLOs, alert routing that actually reaches a human, an on-call rotation, incident response, and postmortems that produce changes. The alerts exist; the paging doesn't.
Own and evolve the IaC - Terraform across AWS, Azure, and GCP, plus the blueprints and per-environment installations that drive it. Bring it tests, policy-as-code, and drift discipline.
Make standing up a new customer environment boring: automate the firewall exceptions, DNS delegations, deploy identities, tag policies, and cert chains that are a runbook today.
Keep the three clouds at parity. A service that exists on Azure but not GCP is a half-shipped service, and we don't get to pick which cloud a customer already bought.
Make HIPAA and SOC 2 properties of the infrastructure rather than projects: controls in code, evidence generated by the pipeline.
Own the observability stack we already run — Grafana, Loki, Mimir, Tempo, Alloy, Langfuse, LiteLLM — and make it answer operational questions about agent behavior: which agent decided what, on whose data, at what cost.
Give the engineers who ship into these tenants a paved road, so deploying doesn't require knowing our blueprint template engine.
What we're looking for
5+ years operating production infrastructure - SRE, platform, or production engineering.
You've operated software in environments you don't control: BYOC, single-tenant, on-prem, or air-gapped. You know what changes when you can't just open the console.
Cloud-agnostic by instinct and comfortable in all three of AWS, Azure, and GCP - networking, IAM, managed Kubernetes (EKS, AKS, GKE), and the operational differences that actually bite. Deep in one is table stakes; we need someone who treats the other two as first-class, not as ports.
Kubernetes in production on managed clusters across multiple providers, and the judgment to know when a cloud-native service is the right call versus a portable one.
Deep Terraform - module design, state, testing, drift, and the judgment to know which blast radius is acceptable.
You've carried a pager and then built the thing that made it quieter: incident command, postmortems, SLOs that changed someone's behavior.
You've worked under a regulated posture — HIPAA, SOC 2, PCI, FedRAMP — where controls had to live in the infrastructure, not a spreadsheet.
Comfortable customer-facing. You will sit in a customer's architecture review and ask their network team for a firewall exception.
Python, Go, or Bash, and the instinct to automate a process the second time you do it.
Genuine curiosity about what you're operating. Agents, not just pods.
Nice to have
You've operated infrastructure through a vendor control plane (Ryvn, Nuon, Replicated, or similar) and have opinions about the tradeoff.
GitOps and progressive delivery across a fleet of tenants running different versions.
Multi-region experience, and a view on where provider-native beats portable in a BYOC product.
Real depth in the Grafana stack — Mimir, Loki, and Tempo at multi-tenant scale.
GPU and inference operations: Ray, SkyPilot, vLLM or SGLang, LLM gateways and cost attribution.
MLOps or research-to-production handoff experience. Not required — you'll be near that boundary either way.
You've thought about what observability means for non-deterministic systems, and what a blast-radius control looks like for something that acts on its own.
The infrastructure patterns for autonomous AI systems are still being written. If you want to be one of the people writing them, let's talk.
About Our Team
As an organization, Percepta is committed to inclusion and belonging, and our people are our greatest investment! We’re working against an incredibly ambitious mission. It won’t be easy, but it will likely be the most fulfilling work of your career.
In addition to competitive compensation, equity, 401(k) matching, comprehensive health and family benefits, and lunch every day, we offer a culture built on autonomy and trust. We strongly believe that the boldest solutions come from teams who dream bigger together, and that means bringing in people from every kind of background, experience, and path. We are an office-first team that values the impact of in-person collaboration, but we focus on outcomes over outputs—believing that we are each responsible for solving the "whole problem" to make our customers win.
If you don't check every box on this description but you care deeply about the problems we're tackling, seeking the truth, and wanting to own outcomes alongside a team that has your back, then we want to hear from you! We know self-doubt keeps great people from applying, so please don't count yourself out before we've had the chance to meet you. Tackling the hardest problems takes a range of perspectives in the room, and we believe that's how we build something truly game-changing.
Our company values are more than just words on a page. They inform how our team shows up every day:
Our Values
Dream bigger: We have the unique privilege of taking on the most ambitious problems and we should chase them with optimism, responsibility, and genuine belief that we can make it happen. We have to embrace the hard things when no one else will.
Heart in the game: What we're doing matters and we have to give a shit. Internally, that means fixing badness when you find it. Externally, it means honoring the trust our customers place in us with their most important problems. This isn’t a 9-5, nor is it a job we’re ever going to monitor your hours. We promise to put work in front of you that matters and in return, we ask you to promise to care. Win for the customer: Everyone is an engineer and the job of an engineer is to deliver outcomes, not outputs. Everything we do—the products we build, the partnerships we launch, the strategy we set—exists to make our customers successful. Delivery is the strategy. Make the call: Organizations are only as strong as the pace at which they make decisions. Everyone at Percepta should feel empowered to commit and shape the ambiguity in front of them. But "make the call" cuts both ways: make the decision and make the phone call. High-agency decision-making only works with high-bandwidth communication and we commit to never operate in silos. Intensity with kindness: We believe in excellence in execution, candor in feedback, ruthlessness in prioritization, and survivalist urgency. We also believe you don't need to be an asshole to deliver on any of this. The trust built through shared kindness and vulnerability is what makes the intensity sustainable. Percepta is proud to be an equal opportunity employer and we evaluate every qualified applicant on their merits and give equal consideration for employment regardless of race, color, creed, religion, sex, gender identity, sexual orientation, national origin, disability, veteran or uniformed service status, age, or any other characteristic protected under applicable law. Percepta provides reasonable accommodations to qualified applicants and employees in accordance with federal and applicable state law.$215k - $350k
We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering...SuggestedWorldwideHome office- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered... ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...SuggestedFull timeWork experience placementLive inWork at officeLocal area
- ...Role Overview: Primary hiring focus is an AI Infrastructure Senior Engineer supporting the build-out of the company's Azure-based technology stack. The role will be the first hire on the infrastructure engineering team and will work under the infrastructure...SuggestedCurrently hiringWork at officeRemote work
$176k - $264k
...Senior AI Software Engineer is required to design and build secure, production-ready AI solutions for one of the top five law firms in... ...resilient AI solutions on Microsoft Azure , collaborating with infrastructure and security teams across cloud, on-premise and hybrid...SuggestedPermanent employmentFull timeH1b3 days per week- ...Type AnnuallyIndustry Law PracticeSelling Points Lead innovative AI platform engineering projects at a forward-thinking organization. Collaborate on cutting-edge Azure AI/ML solutions and cloud infrastructure. Enhance your expertise in MLOps and applied AI technologies....Suggested
$110 - $130 per hour
...diversity in tech, and the best job fit for every candidate we place. Our client, an investment firm, is seeking a Senior AI Platform Engineer to join their team in New York, NY! AIVA is the firm's governed AI execution platform, built to give engineers and...- ...Sr Cloud AI Platform Engineer Contract Hybrid in NYC Must be local. This is not a Full Stack Developer role. We'd like candidates... ...~ Working knowledge of AWS services. ~ Experience using Infrastructure as Code tools, with Terraform preferred. ~ Familiarity...Contract workLocal area
- ...Purpose of Functional Assignment: The Director of AI Platform Engineering provides strategic leadership for the cloud, platform, and deployment infrastructure supporting Artificial Intelligence (AI) across the System. This role ensures AI systems used in clinical workflows...Full time
- ...Postman Fern, the Fern division of Postman, is hiring a Software Engineer in New York. You’ll build APIs, scale AI infrastructure, and design developer experiences for millions of users. Your work will involve real production workloads on Vercel, AWS, and Turbopuffer...
$200k - $230k
SVP, Lead AI/MLOps Infrastructure Engineer - Full Time - HybridWe’re partnering with our client, a fast-growing fintech firm, on a senior-level hire to lead and own the infrastructure behind their AI and machine learning platforms. This is a highly visible, hands-on leadership...Full timeRemote work$180k - $240k
Head of AI Platform Engineering Job Details: Location: New York, NY (Hybrid/Onsite)Employment Type: Full-TimeCompensation: $180-$240k base... ...responsible for building secure, scalable, and resilient AI infrastructure that enables enterprise-wide adoption of AI capabilities....- ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and... ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations...Full timeWork experience placementLive inWork at officeLocal area
$175k - $200k
...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant... ...into production, and our external agent infrastructure (the Agency Connect MCP server, which... ...early partner adoption. Both need a senior engineering owner, and both need to grow into the...Full timeInternship$162k - $215k
...Requirements: We need at least 5 years of experience in infrastructure engineering, with a track record of delivering production-ready systems.... ...platform engineering concepts. We need experience supporting AI and machine learning workloads, including managed compute...Full timeWork at officeWork from homeWorldwideFlexible hours1 day per week$160k - $220k
.... As a global leader in residential wellness and healthcare infrastructure, we create vibrant, purpose-driven communities where housing... ...infrastructure, we want you on our best-in-class team.SUMMARYThe AI Engineer, Platform and Data builds the core that everything else runs...Full time- ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits...
$229.9k - $262.4k
AI Engineer 5 (GenAI Platform Services, Agentic Platform) At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area$229.9k - $262.4k
AI Engineer 5 (Gen AI Platform Services - Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI systems... ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in...Full timePart timeLocal area$250.8k - $286.2k
AI Engineer 5 (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area$229.9k - $262.4k
AI Engineer 5 (GenAI Platform, Agentic Infrastructure) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer...Full timePart timeLocal area- AI Platform Engineer About Tessera Labs Tessera Labs is redefining how enterprises adopt and operationalize Artificial Intelligence. Backed... ...team builds and operates the foundational AI agent infrastructure that lets Tessera run reliably, securely, and consistently...
$230k - $290k
...York( Digital Solutions Group ) - Platform Engineering /Full Time /RemoteAHEAD helps large... ...same enterprises design, build, and run AI agent platforms on top of it. This role... ...delivery leadership attached. You will write infrastructure code, agent code, and the documents that...Full timeWork at officeShift work$155k - $215k
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm's Technology principles... ...of AI across our businesses.This role is for a platform engineering specialist who will help build a firmwide AI Development Platform...Temporary work$144.25k - $256.25k
...(if applicable) + benefitsJob Function: Engineering & ArchitectureSchedule: Full timeShift:... ...technology and data insights.The Enterprise AI Platform organization builds... ...strategiesRetrieval and grounding (RAG) pipelinesLLM infrastructure, inference, and model gatewaysEvaluation...Visa sponsorship$180k - $230k
About the RoleThe Platform Infrastructure team at iCapital plays a critical role in ensuring that... ...and intellectually curious MLOps/DevOps Engineers with deep expertise in machine learning... ...).Enable production workloads for AI/ML and Generative AI systems, including...Full timeWork at officeRemote work- ...Job Title: AI Platform Engineer Location: NYC, NY (Hybrid Model) - 10003 Zip code Energy & Utility domain with experience in Google... ...Platform Engineer: Design and implement AI/ML infrastructure using Vertex AI, Kubeflow, TensorFlow Extended (TFX), and...Local area
- The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data...
- ...Description Job Description Palona’s AI agents operate continuously in... ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...Temporary work
$286.2k - $326.7k
Senior Staff AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems... ...personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in...Full timePart timeLocal areaRemote work$244.7k - $279.2k
Staff AI Engineer - Enterprise Analysis Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
- senior ai engineer New York, NY
- ai developer New York, NY
- ai engineer New York, NY
- ai ml engineer New York, NY
- ai engineer remote New York, NY
- machine learning ai engineer New York, NY
- ai prompt engineer New York, NY
- ai research engineer New York, NY
- remote infrastructure engineer New York, NY
- infrastructure engineer New York, NY






