Staff/Principal DevOps Engineer, AI Inference
$192k - $272kLila Sciences
Staff/Principal DevOps Engineer, AI Inference Cambridge, MA USA The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency. What You'll Be Building GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting What You'll Need to Succeed Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances) Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7 Strong proficiency in Python for automation, tooling, and integration with ML frameworks Bonus Points For Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints Proficiency in Rust or Go for performance-critical infrastructure components SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic Experience with model registries, artifact versioning, and ML supply chain security Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems Compensation We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact. U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program. International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market. Expected Base Salary Range
$192,000 - $272,000 USD
About LILA Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves. LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply. Lila Sciences iscommitted to equal employment opportunityregardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. Information you provide during your application process will be handled in accordance with our Candidate Privacy Policy . A Note to Agencies Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science's internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto. #J-18808-Ljbffr Lila Sciences- Lila Sciences is seeking a Staff/Principal DevOps Engineer for AI Inference in Cambridge, MA. You will design and optimize GPU-accelerated infrastructure for scalable ML model serving, spanning Kubernetes clusters, Terraform/Helm deployments, and multi-region pipelines...Suggested
$118k - $190k
...differences that make us who we are and the work we do possible. Principal DevOps Engineer Boston, MA | Hybrid | Onshape Onshape, a PTC SaaS business... ...more information about PTC’s comprehensive benefits and our AI usage, please visit our Careers Page ( Applications will be...SuggestedFull timeWork at officeLocal areaImmediate startVisa sponsorshipFlexible hours$163.85k - $185k
...enables our customers to develop, deploy and manage responsible, AI-powered applications and experiences with agility and ease.... ...because we believe people power progress. Join us as a Principal DevOps Engineer and help us do what we do best: propelling business forward....SuggestedWork at officeLocal areaWork from homeRelocationHome office$165k - $190k
Maven AGI is an enterprise AI platform founded in July 2023 by executives from HubSpot... ....The RoleWe’re looking for a Senior DevOps Engineer to own and evolve the infrastructure powering... ...GPU resource orchestration, model inference performance tuning, and high-concurrency...Suggested$119k - $221k
...and the customers they serve. We build AI-powered software that keeps everyone in... ...love to meet you. We are looking for a Sr. DevOps Engineer IIto help us build and scale the... ...operating AI/ML infrastructure (model serving, inference, LLM orchestration, or agentic systems)...SuggestedWork at officeLocal areaFlexible hours2 days per week3 days per week$150k - $230k
...impact on its customers, employees, and communities. The Role As a Principal DevOps Engineer for a new product within Veeva, you will be a founding member of a team building our next major AI-driven platform—one that will transform how Life Sciences companies...Work at officeLocal areaRemote workWork from homeFlexible hours- ...Senior Vice President, Principal Full Stack Engineer, Performance Product Engineering About the Company Internationally recognized investment... ...prototype and evaluate new technologies, particularly in the AI and ML space. The ideal candidate will be adept at...
$150k
Boston, MassachusettsHybridDirect Hire$145k - $165kOur client is seeking a Principal DevOps Engineer to join their team in Downtown Boston. This is a full-time, direct-hire role with a hybrid schedule (3 days onsite, 2 days remote). This position is highly hands-on, focused...Full timeRemote work- We are seeking a talented and experienced Principal DevOps Engineer / Site Reliability Engineer (SRE) to lead and drive the DevOps and SRE initiatives for our multi-tenant SaaS platform. The Principal DevOps Engineer will lead the design of our engineering platform, standards...Flexible hours
$142.4k - $213.6k
...to scale, the reliability, security, and maturity of our cloud platform are critical to our success. We are looking for a Principal DevOps Engineer to join our DevOps team and help lead the next stage of our infrastructure, developer platform, and software delivery maturity...Full timeTemporary workFlexible hours$220k - $290k
About the RoleCloudZero is hiring Staff and Principal Software Engineers across our engineering organization. We're not filling a single seat. We're building... ...This team is the spearhead of CloudZero's push into AI cost intelligence. You'll work on the AI telemetry agent...Immediate start- Boston, MAHybridFull Time$120k - $150kWe are seeking a full-time DevOps Engineer to join their team in Chelmsford, MA. This is an opportunity... ...SOC 2 or other compliance framework exposure Experience with AI-assisted development tools (GitHub Copilot, Claude, etc.) Experience...Full time
- ...Meghana GorusuCompany: SRI Tech SolutionsSr. DevOps EngineerBoston, MA (4 days a week onsite)... ...: ContractDescription: The Data Platform Engineering team supports CI/CD and infrastructure... ...with use of advanced features of AI tools: ChatGPT, custom GPTs, and/or other...Hourly pay
- ...MAEmployment Type: Full-TimeJoin a growing engineering team that's modernizing its... ...cloud-native technologies, automation, and DevOps best practices. This full-time opportunity... ...tools SOC 2 or similar compliance experience AI-assisted development tools (GitHub Copilot...Full time
- ...YO AI Labs seeks an experienced Developer & Infrastructure Expert to evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands, configurations, and workflows against real...For contractorsRemote work
$137k - $220k
...you passionate about building secure, compliant cloud platforms that support mission-critical systems? Evolv is seeking a Principal DevOps Engineer to design, operate, and continuously improve cloud infrastructure that meets federal security and compliance requirements....Full timeLocal areaRemote workFlexible hours$175k - $195k
OpenGov is the leader in AI and ERP solutions for local and state governments in the U.S. More than 2,000 cities, counties, state... ...has to be fast, resilient, and built for scale. As a Sr. DevOps Engineer, you'll own the design and delivery of the cloud systems, CI/CD...Contract workWork experience placementWork at officeLocal area- ...customer relationship.The RoleWe're hiring an engineer to own our infrastructure — the AWS... ...including the infrastructure behind our AI agents. And like everyone on this team, you... ...production many times a day with confidenceAttack DevOps problems with AI: use coding agents and...Work at officeLocal areaRemote workWork from home
$250.6k - $362.6k
...performed by a U.S. citizen on U.S. soil.Meet the Team The Platform Engineering organization is responsible for building and operating the... ...data and infrastructure connect and protect organizations in the AI era - and beyond. We’ve been innovating fearlessly for 40 years...Permanent employmentFull timeTemporary workLocal areaFlexible hours- ...institutions. We specialize in leveraging advanced technologies such as AI, cloud, and data-led innovation to help our clients accelerate... ...and do software integration.Required Skill and ExperienceAzure DevOps, Terraform, Harness, IaCPreferred Skill and Experience"Assess...Full timeTemporary workRelocation
$100k - $160k
...people out there, WalkMe is the place for you!We are looking for a DevOps Engineer to join our amazing Cloud Engineering team. We are developing... ...for this specific role.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as...Full timeWork experience placementLocal areaImmediate startRemote work$143k
...thrive.What You'll DoJoin our Data Platform Engineering portfolio, a global team building and... ...powers BCG’s data products, analytics, and AI capabilities across Case Teams, Practice Areas, and Business Systems.As a Principal Database Platform Engineer, you will take...Work at officeLocal area$104k - $169.5k
...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use... ...are looking for a bright, talented DevOps Platform Developer to join a team of highly energized and professional engineers working on mission-critical systems in the red...Full timeWork at officeRemote workVisa sponsorshipWork visa$240k - $330k
...multimodal data mining framework, is the engine that powers this discovery.As a Principal Machine Learning Engineer, you... ...-efficient fine-tuning, optimize inference latency for large-scale retrieval,... ...foundation models, and embodied AI.Hands-on experience pretraining or...Temporary workImmediate start$105.4k - $207.8k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ...you'll doAs a Senior Engineering Management Specialist - DevOps Engineer on the Industry Solutions team, you will be...Local area- ...Senior DevOps EngineerFoley is seeking a Senior DevOps Engineer to help modernize and scale the infrastructure that powers both our core compliance SaaS platform and our next generation of AI-enabled products. This role sits within the Infrastructure Platform team and...Contract workRemote work
$80 - $110 per hour
...is seeking a Senior Platform & Security Engineer to join its Platform Engineering team on... ...minimal supervision, and leveraging modern AI-assisted development tools while maintaining... ...of experience in Platform Engineering, DevOps, SRE, Cloud Engineering, Systems...Hourly payFull timeContract workTemporary workImmediate startFlexible hours$95.2k - $142.8k
...a highly skilled Senior Azure DevSecOps Engineer to design, implement, secure, and optimize... ...that power enterprise software, data, and AI-driven solutions. This role will be a key... ...operations.Strong experience with Azure DevOps, GitHub Actions, CI/CD automation, and release...Full timeSummer workRemote workFlexible hours2 days per week$188k - $258.5k
...reimagining product innovation for the AI era. Our mission is to revolutionize... ...relentlessly focused on impact.THE ROLEAs a Principal Backend Engineer, you’ll be a technical leader... ...excellence.Wear Multiple Hats - Support DevOps, troubleshooting, or other platform needs...- ...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and... ...outcomes.Job SummaryWe are seeking a Principal Engineer Software to lead the technical evolution... ...across large-scale applications.Modern DevOps Practices: Familiarity with...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff/Principal DevOps Engineer, AI Inference. Be the first to apply!
- senior staff systems engineer Cambridge, MA
- assistant engineer Cambridge, MA
- engineering aide Cambridge, MA
- technology administrator Cambridge, MA
- staff engineer Cambridge, MA
- principal developer Cambridge, MA
- senior principal engineer Cambridge, MA
- hotel chief engineer Cambridge, MA
- senior civil engineer project manager Cambridge, MA
- senior chief engineer Cambridge, MA

