Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems Engineer - DevOps& Observability - Senior

$106.9k - $176.5k

EY

Location: Anywhere in Country At EY, we’re all in to shape your future with confidence. We’ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world. We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. These are the systems that ship, run, and make fully visible every AI workload. Within the Hybrid AI Multi-Environment Runtime (HAI), this role advised how AI services and agents are built and deployed, how models execute, how requests are routed to them, how AI assets are catalogued and governed, how consumption is measured and bounded, and how the entire platform is observed across cloud, on-prem, edge, and air-gapped environments. Works with senior engineers to test and develop capabilities. This is a distinct discipline from platform, data, and trust engineering. Where Platform Engineering owns the cluster substrate and its infrastructure automation, this role owns the delivery and runtime surface, including the CI/CD/CV pipelines that ship AI workloads, secure model execution, semantic routing, and model/prompt selection, together with the governance, discovery, cost, and telemetry systems that keep AI workloads shippable, economical, discoverable, and transparent. It sits at the intersection of DevOps, MLOps, FinOps, and observability. This role is ideal for an engineer who is equally comfortable building automated delivery pipelines, operating high-performance inference (GPUs, model servers, sandboxed execution), and building deep observability and cost visibility; who understands that in regulated contexts every AI workload must be delivered repeatably and every AI request must be economically bounded, attributable, and traceable end-to-end. Your Key Responsibilities Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment. Own governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management. Own resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement. Own the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications. Own the OpenTelemetry collection layer, including multi-tenant receiver, exporters and queues (Kafka sink), DCGM exporter for GPU telemetry, processor batching, and dynamic filtering, so every signal is captured and routed reliably. Automate GitOps-based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policy-compliant by default rather than by manual review. Close the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads. Ensure cost and telemetry are identity-stamped and per-tenant, so consumption and behavior are attributable end-to-end, keeping FinOps and observability tied to the workloads that generate the load. Skills And Attributes For Success Strong DevOps expertise: CI/CD/CV pipeline design, GitOps, continuous verification, and progressive/automated release and rollback for production workloads. Deep expertise operating model-serving and inference systems (Ray, vLLM/Triton/NIM) on GPUs at production scale. Deep observability skills: metrics, logs, traces, and OpenTelemetry. FinOps mindset: able to attribute, bound, and optimize AI consumption cost per tenant and workload. Familiarity with model/artifact governance, registries, CVE scanning, and license/lineage tracking. Comfortable operating across cloud, on-prem, edge, and air-gapped environments with consistent runtime and telemetry semantics. Strong communicator able to explain runtime, cost, and observability tradeoffs to engineers, architects, and leadership. To qualify you must have 8+ years in DevOps, MLOps, platform, or observability engineering, with hands-on production ownership of AI or high-throughput services. Strong hands-on DevOps experience, including CI/CD/CV pipelines and GitOps tooling (ArgoCD, Helm, GitHub Actions/GitLab CI, or equivalents) for automated build, test, release, and rollback. Hands-on expertise operating inference/model-serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure. Strong experience with observability stacks (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry. Experience with API gateways and request routing (Envoy or equivalent), including streaming responses. Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rate-limit enforcement. Familiarity with model/artifact registries and supply-chain scanning (Harbor, MLflow, Trivy/SBOM). Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints. Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams. Ideally, you’ll also have Bachelor’s or Master’s degree in Computer Science or related technical field. Experience with LLM evaluation and debugging tooling (LangSmith, Langfuse) and prompt/response quality measurement. Experience with sandboxed/secure execution (gVisor, Firecracker, or microVM isolation) for untrusted or multi-tenant workloads. Familiarity with GPU telemetry (DCGM) and GPU utilization optimization. Experience with lineage and governance contracts (OpenLineage) and AI license management. Exposure to multi-tenant cost attribution and per-tenant SLA/alerting. Exposure to regulated delivery environments (financial services, tax, healthcare, risk). What We Offer You We offer a comprehensive compensation and benefits package where you’ll be rewarded based on your performance and recognized for the value you bring to the business. The base salary range for this job in all geographic locations in the US is $106,900 to $176,500. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $128,400 to $200,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography. In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options. Join us in our team-led and leader-enabled hybrid model. Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year. Under our flexible vacation policy, you’ll decide how much vacation time you need based on your own personal circumstances. You’ll also be granted time off for designated EY Paid Holidays, Winter/Summer breaks, Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being. Learn more. EY | Building a better working world EY is building a better working world by creating new value for clients, people, society and the planet, while building trust in capital markets. Enabled by data, AI and advanced technology, EY teams help clients shape the future with confidence and develop answers for the most pressing issues of today and tomorrow. EY teams work across a full spectrum of services in assurance, consulting, tax, strategy and transactions. Fueled by sector insights, a globally connected, multi-disciplinary network and diverse ecosystem partners, EY teams can provide services in more than 150 countries and territories. EY provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, genetic information, national origin, protected veteran status, disability status, or any other legally protected basis, including arrest and conviction records, in accordance with applicable law. EY is committed to providing reasonable accommodation to qualified individuals with disabilities including veterans with disabilities. If you have a disability and either need assistance applying online or need to request an accommodation during any part of the application process, please call 1-800-EY-HELP3, select Option 2 for candidate related inquiries, then select Option 1 for candidate queries and finally select Option 2 for candidates with an inquiry which will route you to EY’s Talent Shared Services Team (TSS) or email the TSS at View email address on click.appcast.io. EY accepts applications for this position on an on-going basis. EY focuses on high-ethical standards and integrity among its employees and expects all candidates to demonstrate these qualities. #J-18808-Ljbffr EY

Vacancy posted 23 hours ago
Similar jobs that could be interesting for youBased on the AI Systems Engineer - DevOps& Observability - Senior in Boston, MA vacancy
  • $93.95k - $136.74k

     ...people, tech experts, researchers, and systems analysts to advance our mission. As...  ...Mass General Brigham. Job Summary Senior Observability Systems Engineer Join Mass General Brigham’s Digital...  ...combines platform engineering, automation, DevOps practices, and observability... 
    Senior
    Devops
    Full time
    Local area
    Remote work
    Monday to Friday
    Shift work

    Mass General Brigham

    Somerville, MA
    3 days ago
  • $188k - $282k

     ...DescriptionVertex is seeking a Senior Principal AI Engineer to design, build, and...  ...orchestration, evaluation, observability, performance optimization,...  ...experience in applied AI systems. This individual will be comfortable...  ...CI/CD pipelines and DevOps/MLOps practices Secure... 
    Senior
    Devops
    Full time
    Summer work

    Vertex Pharmaceuticals

    Boston, MA
    4 days ago
  • $145k - $230k

     ...person who builds the data and AI platform the rest of the...  ...lives across dozens of SaaS systems and is moved by hand. You'll...  ...agents safely without being cloud engineers.The developer experience for...  ...Kafka, Kinesis, Snowpipe)Data observability and lineage tooling in productionExperience... 
    Senior
    Work at office

    CloudZero

    Boston, MA
    3 days ago
  • $95.2k - $142.8k

     ...seeking a highly skilled Senior Azure DevSecOps Engineer to design, implement...  ...software, data, and AI-driven solutions....  ...teams.Implement observability, monitoring, logging...  ...experience with Azure DevOps, GitHub Actions, CI/...  ...Science, Information Systems, Engineering, or a... 
    Senior
    Devops
    Full time
    Summer work
    Remote work
    Flexible hours
    2 days per week

    Vertex Pharmaceuticals

    Boston, MA
    23 hours ago
  •  ...CDC/ELT integrations from SaaS systems and cloud billing sources...  ...evaluation. Create governed AI Landing Zones across AWS, GCP...  ...at the intersection of data engineering and infrastructure. ~ Deep...  ...Kafka, Kinesis, Snowpipe, data observability and lineage tools, retrieval-... 
    Senior
    Full time
    Work at office

    CloudZero

    Boston, MA
    4 days ago
  •  ...Harvard Business School Publishing Corporation is seeking a Principal DevOps Architect who will drive the DevOps strategy by improving automation and observability throughout the software development lifecycle. Responsibilities include designing a scalable DevOps environment... 
    Senior
    Devops

    Jobleads-US

    Boston, MA
    2 days ago
  • Mass General Brigham is seeking a Senior Observability Systems Engineer to strengthen reliability and performance of critical digital services. You will...  .... This role blends platform engineering, automation, DevOps, and observability expertise, supporting digital systems... 
    Senior
    Devops

    Mass General Brigham

    Somerville, MA
    3 days ago
  • Mass General Brigham Incorporated in Boston is seeking a Senior Observability Systems Engineer to advance enterprise digital services. You’ll deploy...  ...build automation with Terraform, integrating with Azure DevOps and GitHub Enterprise. You will onboard applications, develop... 
    Senior
    Devops

    Mass General Brigham Incorporated

    Somerville, MA
    23 hours ago
  • Mass General Brigham is seeking a Senior Observability Systems Engineer to enhance reliability and visibility across enterprise services. You will work...  ...response. The role emphasizes platform engineering, DevOps practices, and observability expertise with hybrid work... 
    Senior
    Devops

    Worky

    Somerville, MA
    1 day ago
  • $140k - $165k

     ...as well as our leadership in AI maturity and responsible innovation...  ...highly skilled and hands-on Senior AI Engineer to join a small, high-impact...  ..., including multi-agent systems, ML scoring models, and LLM-powered...  ...deployments through Azure DevOps CI/CD and Kubernetes/AKS,... 
    Senior
    Devops
    Work at office
    3 days per week

    InvoiceCloud

    Boston, MA
    23 hours ago
  • Lila Sciences is seeking a Staff/Principal DevOps Engineer for AI Inference in Cambridge, MA. You will design and optimize GPU-accelerated infrastructure for scalable ML model serving, spanning Kubernetes clusters, Terraform/Helm deployments, and multi-region pipelines... 
    Senior
    Devops

    Lila Sciences

    Cambridge, MA
    2 days ago
  •  ...position will report to the MMIS Technical Manager and will serve as the Senior System Analyst for MassHealth’s MMIS application. We are seeking a versatile AWS System Administrator/DevOps Engineer to manage, maintain, and optimize our critical Medicaid Management... 
    Senior
    Devops
    Remote work

    Mindlance

    Quincy, MA
    3 days ago
  • $135k - $325k

     ...OverviewWe are seeking an experienced Engineer to join our Trading Systems team within the Front Office...  ....Explore and apply Generative AI / LLM-based capabilities where...  ...containerization, Kubernetes, DevOps practices, and production observability tools is preferred.Knowledge... 
    Senior
    Devops
    Full time
    Local area

    Arrowstreet Capital

    Boston, MA
    1 day ago
  • $160k - $200k

     ....Tulip, the leader in AI-native frontline operations...  ...Execution System (MES) category.A spinoff...  ...advancements in the realm of Observability & MonitoringYou know...  ...reliability culture across engineering teams. Contributing to...  ...:EngineeringEdge DevOps HardwareWorking At TulipWe... 
    Senior
    Devops
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    2 days ago
  • $125k - $150k

     ...skilled and results-oriented Senior Software Engineer to support the Software...  ...implementation of the design system Implement and promote best front...  ...architecture Leverage AI-powered code assistance tools...  ...systems Experience with Azure DevOps, Team Foundation Server (TFS... 
    Senior
    Devops
    Summer work
    Flexible hours

    Invoice Cloud Inc

    Boston, MA
    23 hours ago
  • VAST Data is seeking a Senior Systems Engineer to join our team in Boston, MA. This position offers the chance to be part of a quickly growing company at the forefront of AI infrastructure. The role involves assisting customers with system installations, developing sales... 
    Senior

    VAST Data

    Boston, MA
    3 days ago
  • $140k - $210.9k

     ...Senior Site Reliability Engineer Federal Reserve Financial Services (FRFS) delivers...  ...come from infrastructure/DevOps backgrounds or software engineering...  ...of distributed production systems. Responsibilities As...  ...Containers, ECR and EKS. Observability - CloudWatch, OpenSearch,... 
    Senior
    Devops
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve Bank of Boston

    Boston, MA
    13 hours ago
  • $140k - $150k

     ...Job Description Job Description Overview SENIOR SYSTEMS ENGINEER LOCATION: Hanscom AFB, MA SALARY RANGE: $140,000-$150,000 annually...  ...performance Agile methodologies, CI/CD, DevSecOps, and DevOps principles Knowledge of systems acquisition and program... 
    Senior
    Devops
    Full time
    For contractors

    Astrion

    Lexington, MA
    2 days ago
  • $173k - $211k

     ...Global AI/ML Engineer Senior Manager - Boston, United States of America Locations : Boston |...  ...Typescript). Proven experience with systems design, design patterns, architectural...  ...Kubernetes, Helm, Terraform). Knowledge of DevOps culture and practices. Experience... 
    Senior
    Devops
    Work at office
    Local area

    Boston Consulting Group

    Boston, MA
    more than 2 months ago
  •  ...Reserve Bank locations As a Senior Engineer of the SRE / Production Operations...  ...to support Engineering, DevOps, and DevSecOps tools,...  ...identify suspected gaps in system architecture and design experiments...  ..., Containers, ECR and EKS. Observability - CloudWatch, OpenSearch, Dynatrace... 
    Senior
    Devops
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    3 days ago
  • $140k - $150k

     ...Job Description Job Description Overview SENIOR SYSTEMS ENGINEER LOCATION: Hanscom AFB, MA SALARY RANGE: $140,000-$150,000 annually...  ...performance Agile methodologies, CI/CD, DevSecOps, and DevOps principles Knowledge of systems acquisition and program... 
    Senior
    Devops
    Full time
    For contractors

    Astrion

    Lexington, MA
    5 days ago
  • $160k - $200k

     ...Overview We are looking for a Sr. Control System Engineer/Site Reliability Engineer (SRE) to...  ...smooth integration/deployment of DevOps tools with custom hardware and quantum...  ...Terraform, or similar). Familiarity with observability tools (e.g., Grafana, Prometheus, ELK... 
    Senior
    Devops
    Local area
    Remote work

    QuEra Computing Inc.

    Boston, MA
    10 hours ago
  • $125k - $140k

     ...Job Description Job Description Overview SENIOR SYSTEMS ENGINEER LOCATION: Hanscom AFB, MA SALARY RANGE: $125,000-$140,000 annually...  ...cost/performance management, Agile, CI/CD, DevSecOps, and DevOps Knowledge of DoD acquisition principles (DoDI 5000.02, 5... 
    Senior
    Devops
    Full time
    For contractors

    Astrion

    Lexington, MA
    14 days ago
  • $140k - $150k

     ...Job Description Job Description Overview SENIOR SYSTEMS ENGINEER LOCATION: Hanscom AFB, MA SALARY RANGE: $140,000-$150,000 annually...  ...cost/performance management, Agile, CI/CD, DevSecOps, and DevOps Knowledge of DoD acquisition principles (DoDI 5000.02, 5... 
    Senior
    Devops
    Full time
    For contractors

    Astrion

    Lexington, MA
    5 days ago
  • CloudZero in Boston is seeking a senior data platform engineer to own the data and AI platform powering the company. You’ll build governed pipelines into Snowflake, model surfaces for analysts and agents, and ensure real identity and cost attribution across AWS, Snowflake... 
    Senior

    Socket.dev

    Boston, MA
    2 days ago
  •  ...Senior AI Engineer, Ares Platform Team: Ares AI Engineering Reports to: Ilir Osmanaj,...  ...Extend the co-evolutionary self-training system that lets Ares learn from its own...  ...Production reliability. Own latency, cost, observability, and failure-mode analysis for agents... 
    Senior
    Full time
    Remote work

    Assail

    Boston, MA
    23 hours ago
  •  ...Description McBride Consulting has an exciting opportunity for a Senior Systems Engineer providing support to the Command, Control, Communications,...  ...cost/performance management, Agile, CI/CD, DevSecOps, and DevOps Knowledge of DoD acquisition principles (DoDI 5000.02, 5... 
    Senior
    Devops
    Full time
    For contractors

    McBride

    Lexington, MA
    14 days ago
  •  ...Job Summary We are seeking a Senior Observability Engineer with strong experience designing, implementing...  ...metrics, logs, and traces to improve system visibility and operational...  ...Collaborate with Platform Engineering, DevOps, SRE, Application Development, and Operations... 
    Senior
    Devops
    Contract work

    2T Consulting

    Boston, MA
    21 days ago
  • $170k - $222.6k

     ...information, click here. Role SummaryJoin Suffolk’s AI Studio in Boston as an AI Systems Engineer, a hybrid role responsible for architecting and building...  .... You will help define how AI is built, deployed, observed, and scaled across Suffolk’s national operations. ResponsibilitiesAI... 
    Temporary work
    For contractors
    Work at office

    Suffolk Construction

    Boston, MA
    4 days ago
  • $90.16k - $167.44k

    AI EngineerJob Posting DescriptionThe AI Engineer will join the AI team supporting the Long-Term Care program in Manulife...  ...MLOps and LLMOps practices.Build observability capabilities, including logging,...  .../AKS, GitHub Actions, or Azure DevOps.Strong SQL skills and experience... 
    Devops
    Full time
    Temporary work
    Local area

    John Hancock

    Boston, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems Engineer - DevOps& Observability - Senior. Be the first to apply!