Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer — AI Performance & Reliability

AMD

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE:We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.THE PERSON:Profile and optimize AI model training and inference workloads.Improve model throughput, latency, memory efficiency, scalability, and reliability.Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.Develop performance tooling, benchmarks, automation, and observability systems.Investigate and resolve complex production issues affecting AI workloads.Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.Translate customer feedback into product and infrastructure improvements.Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.Document performance findings, technical recommendations, and best practices.KEY RESPONSIBILITIES:Strong software engineering skills and experience building production-quality systems.Experience working with AI infrastructure for model training, inference, or both.Demonstrated experience profiling and optimizing machine learning models or AI workloads.Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication.Proficiency in languages such as Python, C++, or similar systems-oriented programming languages.Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack.Clear written and verbal communication skills.A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.PREFERRED EXPERIENCE:Experience optimizing large language models, diffusion models, or recommendation models.Experience with GPU, accelerator, or distributed computing environments.Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.Experience operating AI systems in production environments.Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role.Experience designing benchmarks and conducting systematic performance analysis.ACADEMIC CREDENTIALS:A PhD (or a master’s degree with equivalent experience) in artificial intelligence, machine learning, computer science, or a related field. LOCATION:San Jose, CA or Bellevue, WA preferred (Hybrid). Other US locations may be considered.#LI-MV1#HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer — AI Performance & Reliability in San Jose, CA vacancy
  • $250.6k - $362.6k

     ...Security team builds software and cloud product...  ...team connects security engineering, security operations...  ....Your ImpactAs a Principal Software Engineer, you...  ...distributed systems, reliability, and performance. You will mentor...  ...organizations in the AI era - and beyond. We... 
    Performance
    Full time
    Temporary work
    Work experience placement
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $250.6k - $362.6k

     ...TeamSplunk’s Platform Engineering team is evolving the core platform to run reliably on Kubernetes across cloud...  ...years of professional software engineering experience...  ...experience using AI-assisted engineering tools...  ...on sales plans earn performance-based incentive pay on... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    9 hours ago
  • $250.6k - $362.6k

     ...(CVIS) establishes and proves the performance of large AI clusters before they are handed over...  ...environments. Your Impact  As a Principal Software Engineer, you will be the hands-on architect...  ...capacity, delivery cycle time, reliability, or user adoption—and coordinating... 
    Performance
    Full time
    Temporary work
    Work experience placement
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 hours ago
  • $135.2k - $306.4k

     ...are hoping to enhance engineering efficiency by concentrating...  ...systems with high performance that can be adopted by...  ...that will ensure the reliability of databases being used...  ...As a Senior Principal Engineer, you will lead...  ...saving care. And with AI embedded across our products... 
    Performance
    Temporary work
    Worldwide
    Flexible hours

    Oracle

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the next era...  ...firmware and software architecture and design...  ...passion for building reliable, debuggable, and scalable...  ...Mentor architects and engineering teams to grow them into...  ...architecture for scalable and performant edge systems,... 
    Performance
    Shift work

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $221.2k - $387.1k

     ...It all started when engineer Fred Luddy wrote code...  ...Today, ServiceNow is the AI control tower for business...  ...the quality and reliability of our security and risk...  ...scalability, reliability, performance, and maintainability...  ...foundation in software architecture, distributed... 
    Performance
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    28 days ago
  • $221.2k - $387.1k

     ...all started when engineer Fred Luddy wrote code...  ...ServiceNow is the AI control tower for...  ...quality, and software delivery efficiency...  ...You'll Do As a Principal Software Engineer,...  ...related to scalability, reliability, observability, security, and performance across customer-... 
    Performance
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    18 days ago
  • $221.2k - $387.1k

     ...It all started when engineer Fred Luddy wrote code...  ...Today, ServiceNow is the AI control tower for...  ...looking for an experienced Principal Software Engineer to own the...  ...~ Own production reliability for the platform — participate...  ...— schema design, performance tuning, and trade-... 
    Performance
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    17 days ago
  •  ...searching for a passionate Principal Software Engineer to join our engineering...  ...delivers a fast, seamless, and reliable experience for customers....  ...unit testing, automation, performance optimization, and observability...  ...from top operators. AI-First Skill Building: Get... 
    Performance
    Local area

    CSC Generation

    San Jose, CA
    3 days ago
  • $114.6k - $234.6k

     ...industry innovations to life-saving care. And with AI embedded across our products and services, we...  ...retrieval, storage, and processing. -Design performance and load testing.System Design & Architecture - System Reliability Design:-Build and design fault-tolerant components... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    2 hours ago
  • $249k

     ...for travelers everywhere.Principal Software Development Engineer Our Technology Team partners...  ...without sacrificing reliability or skyrocketing our cloud...  ...: While we are embracing AI tools, your job is to build...  ...including capacity planning, performance optimization, and robust... 
    Performance
    Full time
    Weekend work

    Expedia

    San Jose, CA
    9 hours ago
  • $210k - $275k

     ...solutions that accelerate AI data centers to...  ...Faster AI. Today’s AI performance is frequently limited...  ...across silicon, packaging, software, and systems to...  ...power costs and improving reliability. The company’s solutions...  ...and has a world‑class engineering team with decades of... 
    Performance

    Eridu Corporation

    Saratoga, CA
    4 days ago
  • $210k - $275k

     ...Staff - Principal, Software Engineer, SDK Eridu is a Silicon Valley-based hardware...  ...infrastructure solutions that accelerate AI data centers to deliver Faster AI. Today’s AI performance is frequently limited by...  ...power costs and improving reliability. The company’s solutions... 
    Performance

    Eridu

    Saratoga, CA
    4 days ago
  • $240k - $379.5k

     ...Platform and collaborate with engineering to build the platform. This...  ...and knowledge of enterprise AI architectures will make you...  ...efficiency. Ensure system reliability through monitoring,...  ...Artificial Intelligence, High-Performance Computing, and Visualization... 
    Performance
    Work experience placement

    NVIDIA Gruppe

    Santa Clara, CA
    3 days ago
  • $170k - $277k

     ..., and Inclusion. We weave AI into the fabric of everything...  ...We are seeking a Senior Principal Software Engineer who is first and foremost...  ...excellence, and platform reliability. If you thrive on solving...  ...for maintainability, performance, security, and operational... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks, Inc.

    Santa Clara, CA
    2 days ago
  • $96.8k - $306.4k

     ...with significant impact on reliability, performance, and compliance across the...  ...life-saving care. And with AI embedded across our products...  ...Key ResponsibilitiesPlatform Software Development:Set group-wide...  ...lifecycle; coaches engineers across teams or units to drive... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    2 days ago
  • $240k - $250k

     ...Description Saviynt's AI-powered identity...  ...excellence, reliability, disaster recovery...  ...AI & Agentic Engineering Apply AI-assisted...  ...best practices for software delivery,...  ...BRING ~1+ years of Principal-level of platform...  ...and organizational performance.  Saviynt is an... 
    Performance

    Saviynt

    Milpitas, CA
    a month ago
  • $240k - $250k

     ...Description Saviynt's AI-powered identity...  .... Optimize agent performance, scalability, reliability, and resource utilization...  ...and mentor engineers building endpoint platform...  ...AI-assisted software development throughout...  ...BRING ~1+ years of Principal-level of systems software... 
    Performance
    Local area

    Saviynt

    Milpitas, CA
    a month ago
  •  ...globally for innovation, performance and quality....  ...Description We are hiring a Principal Engineer to serve as an...  ...for Nexus, Enterprise AI platform for engineering...  ...scales securely and reliably as adoption grows...  ...years of professional software engineering experience... 
    Performance
    Temporary work
    Remote work
    Flexible hours
    Shift work

    SanDisk

    Milpitas, CA
    4 days ago
  • $240.1k - $420.2k

     ...Principal Data Platform Software Engineer (RaptorDB) Full-time Employee Type: Regular...  ...Today, ServiceNow is the AI control tower for business...  ...critical solutions that enable reliable service for millions of...  ...standards for uptime, performance, and resilience. Role... 
    Performance
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    SmartRecruiters, Inc.

    Santa Clara, CA
    1 day ago
  • $207k - $300k

     ...members to enhance system reliability and efficiency....  ...reliability, scalability, and performance of SU services, often...  ...of experience with software development in one or...  ...as a Site Reliability Engineer.3 years of experience...  ...Experience in Generative AI, Generative AI Agent,... 
    Performance

    Google

    San Jose, CA
    2 hours ago
  • $156.4k - $253k

     ...Integrity, and Inclusion. We weave AI into the fabric of everything...  ...the Layer-7 Security Software team, we are responsible for...  ...Identification and Content Inspection Engine runs on Hardware, Virtualized...  ...terms of functionality and performance, working on device identity... 
    Performance
    Full time
    Work at office

    Palo Alto Networks

    San Jose, CA
    4 days ago
  • $165.6k - $296.4k

     ...than 25%Profession: Software...  ...Role We’re building AI‑first growth and...  ...designing foundational engineering systems (instrumentation...  ...confidence. As a Principal Growth Engineer in...  ...that improve reliability and decision quality...  ...quality (reliability, performance, operability,... 
    Performance
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    1 day ago
  • $143k - $286k

     ...Home OfficeRole summary: As a Principal Software Engineer, you will lead the design...  ...solutions that support AI/ML integration and real-time...  ...business needs into innovative, reliable software solutions that...  ...building highly available, high-performance, redundant, and scalable... 
    Performance
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    2 hours ago
  • $119.8k - $234.7k

     ...: Less than 25%Profession: Software EngineeringDiscipline: Software...  ...the Role We’re building AI‑first engineering systems that power growth...  ...Proficiency in designing scalable, reliable systems that support rapid...  ...Ability to reason about performance, reliability, and... 
    Performance
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    1 day ago
  • $132.6k - $214.5k

     ...Integrity, and Inclusion. We weave AI into the fabric of...  ...collaborate closely with our engineering teams to develop innovative...  ...insights into our systems’ performance and health. As a Senior Staff...  ...the product and ensure the reliability and availability of our services... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $161k - $299k

     ...verification technology. As a Senior Engineer in the System Verification...  ...complexity of future AI and hyperscale chip designs....  ...looking for an experienced C/C++ software engineer to join the Xcelium...  ...bottleneck analysis and implement performance optimizations in C/C++ to... 
    Performance

    Cadence Design Systems

    San Jose, CA
    1 day ago
  • $183.6k - $297k

     ...Integrity, and Inclusion. We weave AI into the fabric of everything...  ...life. We are looking for an Engineering Manager to lead the Explicit...  ...applications - with scale, performance, and zero-trust principles at...  ...of experience managing a software engineering team in a large enterprise... 
    Performance
    Full time
    Work at office

    Palo Alto Networks, Inc.

    Santa Clara, CA
    1 day ago
  • $161k - $299k

     ...technology. Why Cadence? — Design for AI. AI for Design. Cadence sits at the intersection...  ...of semiconductor innovation and high‑performance computing. Through Design for AI, our...  ...team, you will build high‑performance software engines and algorithms behind next‑generation... 
    Performance

    Cadence Design Systems

    San Jose, CA
    23 hours ago
  • $175k - $265k

     ...potential of generative AI to power the...  ...are at the forefront of software and hardware innovation...  ...infrastructure layer that every engineering team and customer...  ...team, responsible for reliability, automation, and observability...  ...platform services.Perform hands-on... 
    Performance

    d-Matrix

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer — AI Performance & Reliability. Be the first to apply!