Staff Software Development Engineer - Enterprise AI Infrastructure - #4898
$169k - $224kGRAIL
Job Description
Job Description
Our mission is to detect cancer early, when it can be cured. We are working to change the trajectory of cancer mortality and bring stakeholders together to adopt innovative, safe, and effective technologies that can transform cancer care.
We are a healthcare company, pioneering new technologies to advance early cancer detection. We have built a multi-disciplinary organization of scientists, engineers, and physicians and we are using the power of next-generation sequencing (NGS), population-scale clinical studies, and state-of-the-art computer science and data science to overcome one of medicine’s greatest challenges.
GRAIL is headquartered in the bay area of California, with locations in Washington, D.C., North Carolina, and the United Kingdom. It is supported by leading global investors and pharmaceutical, technology, and healthcare companies.
For more information, please visit grail.com
The Staff Software Development Engineer - Enterprise AI Infrastructure is a senior technical role responsible for leading the design, development, and scaling of a centralized, highly governed enterprise AI platform. This position serves as a technical expert focused on AWS and Kubernetes-based (EKS) AI infrastructure, agentic development, and multi-agent orchestration operating within a regulated environment. The Staff Engineer partners closely with cross-functional stakeholders across Software Engineering, Data Science, Security, Regulatory Affairs, and Product to build a unified control plane that securely connects large language models with enterprise tools and company knowledge.
This role is expected to drive technical excellence in cloud infrastructure, container orchestration, AI governance, identity-scoped integrations, and agentic workflows while mentoring engineering teams and advancing the organization's enterprise AI strategy.
This role is based in Sunnyvale, California in our new headquarters. We will move in September so you may potentially visit our current location in Menlo Park, CA for interviews. We will also consider candidates in our Durham, NC office. We offer a flexible work arrangement, with the ability to work from GRAIL's office or from home. Our current flexible work arrangement policy requires that a minimum of 60%, or 24 hours, of your total work week be on-site. Your specific schedule, determined in collaboration with your manager, will align with team and business needs and could exceed the 60% requirement for the site. At our Sunnyvale and Durham campuses, Tuesdays and Thursdays are the key days where we encourage on-site presence to engage in events and on-site activities.
Responsibilities- Lead the end-to-end design, development, deployment, and monitoring of a scalable, governed enterprise AI platform leveraging Amazon EKS and AWS native services (e.g., Bedrock, OpenSearch Serverless, KMS, VPC).
- Design and implement agentic AI workflows, specialized autonomous agents, and multi-agent systems using advanced LLM orchestration techniques and agent frameworks.
- Architect and manage secure integrations using the Model Context Protocol (MCP) to connect the AI platform with internal systems, vector databases, and third-party SaaS applications (e.g., Google Workspace, Slack).
- Build and enforce strict identity, authorization, and zero-trust token brokering flows leveraging Okta, Auth0, and custom JWT authorizers to ensure secure, least-privilege tool execution.
- Implement deterministic policy controls (e.g., Cedar policy engine) to enforce role-based access, approval gates, and human-in-the-loop checks at the API gateway level.
- Develop and maintain highly isolated, scalable containerized runtime environments (e.g., Kubernetes pods on Amazon EKS) for secure AI model execution, tool usage, and knowledge retrieval.
- Establish and maintain comprehensive audit trails and observability for all AI interactions, utilizing AWS CloudTrail and GenAI observability tools (e.g., OpenTelemetry) to track cost, latency, and tool calls.
- Collaborate with Product Management, Security, Regulatory, and business stakeholders to translate enterprise requirements into scalable, compliant AI infrastructure solutions.
- Troubleshoot and resolve complex technical issues involving cloud infrastructure, Kubernetes networking, network isolation (PrivateLink), and agentic workflows.
- Contribute to technology roadmaps, AI infrastructure strategy, and long-term platform evolution initiatives.
- Mentor engineers, software developers, and technical teams while promoting engineering excellence, infrastructure-as-code (IaC) best practices, and continuous improvement.
- Partner with Quality, Regulatory, Privacy, Security, and Compliance functions to ensure software and AI systems operate in accordance with applicable regulatory requirements and company policies.
As our organization continues to evolve and grow, this role may require flexibility in responsibilities and duties. Employees should expect that their role may expand, shift, or be modified to meet changing business needs, strategic priorities, and organizational objectives.
This may include:
- Taking on additional responsibilities.
- Participating in cross-functional projects and initiatives.
- Adapting to new technologies, AI methodologies, software frameworks, processes, or engineering practices.
- Supporting other departments or teams during periods of high demand.
- Contributing to special projects or temporary assignments as needed.
These job duties are a summary of the primary duties and responsibilities of the position and are not intended to be a comprehensive or all-inclusive listing of duties. Contents are subject to change at the Company's discretion.
Required Qualifications- Bachelor's degree or equivalent in Computer Science, Software Engineering, Artificial Intelligence, Cloud Computing, or related field; Master's or PhD preferred.
- 8-12 years of relevant software development and cloud infrastructure experience with demonstrated technical leadership.
- Deep expertise in AWS cloud architecture and container orchestration, specifically with Amazon EKS, Kubernetes networking, network isolation (VPC, PrivateLink), IAM, KMS, and GenAI services (e.g., AWS Bedrock).
- Proven experience in agentic AI development, building autonomous agents, and orchestrating LLM tool-calling workflows using frameworks like LangChain, LangGraph, AutoGen, or Claude Agent SDK.
- Hands-on experience implementing the Model Context Protocol (MCP) or building robust, governed API/tool integrations for LLMs.
- Strong background in identity and access management (IAM), OAuth, JWT, and integrating with enterprise IdPs (Okta, Auth0) for scoped, token-based authorization.
- Advanced proficiency in programming languages such as Python, TypeScript, or Go, and infrastructure-as-code tools (Terraform, AWS CDK).
- Experience with vector databases, RAG (Retrieval-Augmented Generation) architectures, and row-level access controls (e.g., OpenSearch, FAISS, pgvector).
- Proficiency with CI/CD pipelines, MLOps practices, Kubernetes ecosystem tools (e.g., Helm), containerization, and modern observability stacks.
- Demonstrated level of knowledge regarding applicable regulatory standards commensurate with the position's complexity and scope, contributing to organizational regulatory compliance. Minimal applicable standards for this position include:
- Cybersecurity principles, tools, and control frameworks (e.g., ISO 27001, NIST, SOC 2, HIPAA)
- Operations within the regulated medical device environment (e.g., IVDD, IVDR, FDA 21 CFR 800 series, FDA 21 CFR Part 11)
- AI governance, software validation, data integrity, and risk management principles applicable to regulated environments
- Deep expertise in cloud infrastructure, containerized environments, agentic artificial intelligence, and secure distributed system design.
- Exceptional problem-solving and analytical skills with the ability to address ambiguous, high-impact technical challenges in AI orchestration and Kubernetes scaling.
- Strong leadership and influence skills, capable of driving alignment across engineering, security, regulatory, and business stakeholders.
- Excellent communication skills with the ability to explain complex LLM behaviors, infrastructure architectures, and security boundaries to technical and non-technical audiences.
- Proven mentoring and coaching capabilities that elevate cloud engineering and AI talent.
- Strong understanding of AI safety, prompt injection defenses, secure tool execution, and deterministic policy enforcement.
- Strategic thinking with the ability to balance long-term enterprise AI platform vision with near-term business delivery.
- High adaptability and intellectual curiosity regarding emerging agentic AI frameworks, MCP specifications, and cloud computing trends.
- Standard office or hybrid work environment depending on company policy.
- Frequent use of software development tools, AI/ML platforms, cloud infrastructure, data engineering tools, and collaboration systems.
- May require extended hours during major project deadlines, AI model deployments, production incidents, regulatory audits, or strategic initiatives.
- Operates with significant independence and responsibility and is expected to provide leadership on complex software, AI, technical, and organizational decisions.
The expected, full-time, annual base pay scale for this position is $169k-224k in Sunnyvale, CA and $147K-$195K in Durham NC.
This role may be eligible for other forms of compensation, including an annual bonus and/or incentives, subject to the terms of the applicable plans and Company discretion. This range reflects a good-faith estimate of the range that the Company reasonably expects to pay for the position upon hire; the actual compensation offered may vary depending on factors such as the candidate’s qualifications. Employees in this role are also eligible for GRAIL’s comprehensive and competitive benefits package, offered in accordance with our applicable plans and policies. This package currently includes flexible time-off or vacation; a 401(k) retirement plan with employer match; medical, dental, and vision coverage; and carefully selected mindfulness programs.
GRAIL is an equal employment opportunity employer, and we are committed to building a workplace where every individual can thrive, contribute, and grow. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, gender, gender identity, sexual orientation, age, disability, status as a protected veteran, , or any other class or characteristic protected by applicable federal, state, and local laws. Additionally, GRAIL will consider for employment qualified applicants with arrest and conviction records in a manner consistent with applicable law and provide reasonable accommodations to qualified individuals with disabilities. Please contact us at View email address on us.fitly.work if you require an accommodation to apply for an open position.
GRAIL maintains a drug-free workplace. We welcome job-seekers from all backgrounds to join us!
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
- ...world's largest AI chip, 56 times larger... ...labs, global enterprises, and cutting-edge... ...scale.Partner with development teams to design... ...and optimize cloud infrastructure supporting CI... ...developer velocity and engineering productivity.... ...professional experience in software engineering,...Suggested
$207k - $300k
...Warsaw.Develop junior engineers on the team.... ...testing, and launching software products.5 years... ...large-scale infrastructure, distributed systems... ...through AI by combining cutting... ...tools drive rapid development, rooted in a culture... ...the frontier of enterprise and driving the evolution...Suggested$207k - $301k
...and test system development code to... ...designing robust infrastructure solutions that... ...and launching software products.5 years... ...degree or PhD in Engineering, Computer... ...workflows or AI transformations... ...destination for Enterprise customers to discover... ....As a Staff Software Engineer...Suggested$207k - $301k
...designers, and infrastructure teams to ship... ...Chiliagon and Agentic Development Kit (ADK) to... ...Lifecycle AI Tools).Mentor... ...to other engineers and scientists... ...experience in software development.5... ...market leader.As a Staff Software Engineer... ...for our enterprise and developer...Suggested$215k - $352k
...in Mountain View, CA.The Enterprise Infrastructure Engineering (EIE) team builds and operates... ...like LangSmith, AI integrations, and enterprise... ...(Sr. Managers, Managers, Staff+ ICs); foster a culture of... ...experience in infrastructure/software engineering with 8+ years...SuggestedFor contractorsWork at officeRemote workFlexible hours$207k - $300k
...maintain, and enhance large scale software solutions.Provide technical leadership... ...and coach a distributed team of engineers.Lead the design and implementation... ...specialized ML areas, optimize ML infrastructure, and guide the development of model optimization and data...$193.93k - $352.29k
...and profound opportunity for AI to drive positive change in the... ...investors. About the Role Our software team is growing, and we are looking for talented engineers to join us and be instrumental... ..., Simulation, and Technical Infrastructure. Data Platform: The Data...Immediate startFlexible hours$193.93k - $352.29k
...and profound opportunity for AI to drive positive change in the... ...Role The Autonomy ML Infrastructure team is responsible for building... .... Work with autonomy engineers to optimize, validate, and deploy... ...Write robust, high quality software to increase our confidence in...Work experience placementImmediate startFlexible hours- ...computing and Confidential AI for hybrid and multicloud environments... ...across clouds, on-premises infrastructure, and devices. Our... ...security. The Role Staff Software Engineer (Rust) - Confidential... ...confidential computing / TEE development (Intel SGX, Intel TDX, AMD...Full timeH1b
$300 per month
...only vertically integrated AI infrastructure company built from the ground... ...Data Center Infrastructure Engineering (DCIE) team is fundamental... ...skilled and motivated Software Engineer to join Crusoe’s Data... ...position is focused on the development of software for the management...Temporary work$193.93k - $352.29k
...and profound opportunity for AI to drive positive change in the... ...leading investors.About the RoleOur software team is growing, and we are looking for talented engineers to join us and be instrumental... ...and tracing tools and infrastructure (perf, eBPF, Perfetto, pprof,...Immediate startFlexible hours$198k - $326k
...Join us in building the AI Governance Platform... ...critical path between model development and production... ...repeatable, and enforceable engineering capabilities.As a Sr. Staff Software Engineer, you will help... ...generation of AI Governance infrastructure. You will solve complex...For contractorsWork at officeFlexible hours- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of... ...from AI researchers to enterprises and hyperscalers. Lambda... ...Lambda Infrastructure Engineering organization forges the... ...are seeking a seasoned Staff Storage Software Engineer with deep...Work at officeLocal areaWork from homeFlexible hours
$240k - $265k
...artificial intelligence (AI) powered technology... ...commercial self-driving software to develop, test and deploy... ...looking for a Senior or Staff Software Engineer to build infrastructure, tools, and systems that... ...behavior and motion planning development. You will work closely...Visa sponsorship$207k - $300k
...a distributed team of engineers.Facilitate alignment and... ...enhance large-scale software solutions.Minimum qualifications... ...large-scale infrastructure, distributed systems or... ...forward.The AI and Infrastructure team... ...innovations, empowering the development of our cutting-edge AI...Worldwide$198k - $326k
...scaling LinkedIn's AI model training, feature engineering and serving with... ...data infra, compute software, and hardware to... ...queries.Model Training Infrastructure: As an engineer on... ...in and guide the development of containerized... ...at scale.As a Sr. Staff Software Engineer,...For contractorsWork at officeFlexible hours$175k - $287k
...privately managed compute infrastructures in the world outside the public... ...cloud providers. As a Staff Software Engineer on the Compute... ...petabytes daily, and the AI/ML infrastructure driving... ...experience in software design, development, and algorithm-related solutions...For contractorsWork experience placementWork at officeFlexible hours$193.93k - $352.29k
...profound opportunity for AI to drive positive... ...is not fungible is the infrastructure that decides whether an... ...autonomously inside Nuro's own engineering organization, under... ...You ~5+ years of software engineering experience... ...practical experience. Staff-level candidates...Immediate startFlexible hours$231k - $378k
...team.We are seeking a Principal Staff Software Engineer to join LinkedIn’s Physical Infrastructure organization. As the technical... ...organization. You will drive the development of reliable, scalable, and... ...requirements such as large-scale AI clusters, new hardware architectures...For contractorsWork at officeImmediate startFlexible hours$198k - $326k
...part of our world-class software engineering team, you will take... ...the next-generation infrastructure and platforms for LinkedIn... ..., best-in-class AI/ML infrastructure, Kubernetes... ...our company.As a Sr. Staff Software Engineer,... ...in software design, development, and algorithm...For contractorsWork at officeFlexible hours$207k - $300k
...a distributed team of engineers.Facilitate alignment and... ...enhance large scale software solutions.Minimum qualifications... ...in software development.5 years of experience... ...developing large-scale infrastructure, distributed systems or... ...solutions.The AI and Infrastructure team...Worldwide$207k - $301k
...a distributed team of engineers.Facilitate alignment and... ...enhance large scale software solutions.Minimum qualifications... ...in software development.5 years of experience... ...developing large-scale infrastructure, distributed systems or... ...solutions.The AI and Infrastructure team...Worldwide$100k
...industry on cutting-edge AI technology,... ...to unify innovations in software models, compilers, platforms... ...looking for a hands-on Staff Software Engineer with a platform infrastructure / Site Reliability Engineering... ...The role spans backend development, service integrations,...Permanent employment$207k - $300k
...and coach a distributed team of engineers.Scale and performance tuning... ..., maintain and improve switch software.Work on projects enabling AI networking infrastructure, enabling high performing fabrics... ...Knowledge of Linux user space development and multi-threading...Worldwide$198k - $326k
...business needs of the team. LinkedIn’s AI Infrastructure organization is responsible for... ...frameworks.We are looking for a Senior Staff Software Engineer with deep expertise at the... ...systems.ResponsibilitiesLead the design, development, and optimization of LinkedIn’s large...For contractorsWork at officeFlexible hours$262k - $364k
...coach a distributed engineering team, fostering... ...deploy scalable software and advanced serving... ..., storage, and AI/ML, staying ahead... ...Machine Learning Infrastructure.Google's software... ...GDS).As a Senior Staff Software Engineer... ..., empowering the development of our cutting-...Remote workWorldwide$193.93k - $352.29k
...and profound opportunity for AI to drive positive change in the... ...a scalable and reliable data infrastructure. This infrastructure is... ...collaborates closely with system engineers to thoroughly validate the autonomous... ...practices across broader software organizationA bachelor's...Immediate startFlexible hours$207k - $300k
...adoption of next-generation AI/ML techniques to... ...and mid-level engineers, fostering a culture... ...years of experience in software development.5 years of experience... ...decision making), ML infrastructure, or specialization in... ...technology forward.As a Staff Software Engineer in...Immediate start$262k - $365k
...a distributed team of engineers.Facilitate alignment and... ...enhance large-scale software solutions.Build next generation... ...generative AI tools or LLM interfaces... ...engineering is a critical infrastructure service for Google. In... ..., empowering the development of our cutting-edge AI...Worldwide$226k - $369k
...approval. We are seeking a Principal Staff Software Engineer to join our organization. The team... ...and driving next-generation infrastructure that powers AI-first unified developer platforms... ...design standards across the software development lifecycleSet infrastructure architecture...For contractorsWork at officeRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Development Engineer - Enterprise AI Infrastructure - #4898. Be the first to apply!
- software engineer internship Sunnyvale, CA
- software development engineer aws Sunnyvale, CA
- software developer internship no experience Sunnyvale, CA
- real time software engineer Sunnyvale, CA
- financial software developer Sunnyvale, CA
- part time software developer Sunnyvale, CA
- graduate software developer Sunnyvale, CA
- software engineer travel Sunnyvale, CA
- experienced software developer Sunnyvale, CA
- remote entry level software developer Sunnyvale, CA


