Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer — AI Performance & Reliability

Jobleads-US

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.

Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.

This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.

You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.

THE PERSON:

  • Profile and optimize AI model training and inference workloads.
  • Improve model throughput, latency, memory efficiency, scalability, and reliability.
  • Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
  • Develop performance tooling, benchmarks, automation, and observability systems.
  • Investigate and resolve complex production issues affecting AI workloads.
  • Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.
  • Translate customer feedback into product and infrastructure improvements.
  • Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.
  • Document performance findings, technical recommendations, and best practices.

KEY RESPONSIBILITIES:

  • Strong software engineering skills and experience building production-quality systems.
  • Experience working with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication.
  • Proficiency in languages such as Python, C++, or similar systems-oriented programming languages.
  • Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack.
  • Clear written and verbal communication skills.
  • A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.

PREFERRED EXPERIENCE:

  • Experience optimizing large language models, diffusion models, or recommendation models.
  • Experience with GPU, accelerator, or distributed computing environments.
  • Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.
  • Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.
  • Experience operating AI systems in production environments.
  • Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role.
  • Experience designing benchmarks and conducting systematic performance analysis.

ACADEMIC CREDENTIALS:

  • A PhD (or a master’s degree with equivalent experience) in artificial intelligence, machine learning, computer science, or a related field.

LOCATION:

San Jose, CA or Bellevue, WA preferred (Hybrid). Other US locations may be considered.

#LI-MV1

#HYBRID

  • Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer — AI Performance & Reliability in San Jose, CA vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team...  ...focal point for fleet-scale reliability, working directly with...  ...Artificial Intelligence, High-Performance Computing and...  ...existing vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...delivery across system software, drivers, and CUDA to...  ...continuously available and reliable.What you’ll be doing:...  .../platform layers, and performance counter/trace...  ...integrate with existing ML/AI workflows (e.g., PyTorch...  ...direction for an engineering team; mentor engineers... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the next era...  ...firmware and software architecture and design...  ...passion for building reliable, debuggable, and scalable...  ...Mentor architects and engineering teams to grow them into...  ...architecture for scalable and performant edge systems,... 
    Performance
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as...  ...to ensure they can reliably manage, update, and operate...  ...secure boot, attestation), and performance — and champion those priorities...  ...vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $248k - $391k

     ...scalable simulation, AI, and thorough...  ...systems that help engineers develop, test,...  ...scenarios. In this Principal-level individual...  ...parameter adaptation, software interfaces,...  ...safety, simulation reliability, synthetic data,...  ...dashboards, or high-performance numerical... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...learning ignited modern AI—the next era of...  ...assistants and engineering-productivity...  ...company. Now we need a principal-level, hands-on...  ...obsesses over reliability, polish, and user...  ...Improve reliability, performance, observability,...  ...like mature software, not prototypes.Build... 
    Performance
    Full time
    Live in

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the next era...  ...maps improve driving performance, safety, and coverage....  ...with a diverse team of engineers in mapping, perception...  ...fleet data into reliable map products used in self...  ...building production-quality software systems.Solid... 
    Performance
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $221.2k - $387.1k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  ..., ServiceNow is the AI control tower for...  ...the quality and reliability of our security and risk...  ...scalability, reliability, performance, and maintainability....  ...technical foundation in software architecture,... 
    Performance
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  • $220k - $250k

     ...United StatesProducts - Engineering /Fulltime /HybridOver...  ...of position: Principal Software EngineerPosition type...  ...networking, generative AI, and autonomous agentic...  ...be scalable, secure, reliable, observable, adaptable...  ...networks by building high-performance, real-time systems... 
    Performance
    Full time
    H1b
    Local area
    Work from home
    Work visa
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    4 days ago
  • $249k - $348.5k

     ...Principal Software Development Engineer Our Technology Team partners with teams across...  ...by default. Hardened Reliability & Observability: Set SRE...  ...direction. Familiarity with AI‑driven systems and...  ...including capacity planning, performance optimization, and robust... 
    Performance
    Flexible hours

    Traveltechessentialist

    San Jose, CA
    2 days ago
  • $250.6k - $362.6k

     ...Security team builds software and cloud product...  ...team connects security engineering, security operations...  ...Your Impact As a Principal Software Engineer,...  ...distributed systems, reliability, and performance. You will mentor engineers...  ...in the AI era – and beyond. We... 
    Performance
    Full time
    Temporary work
    Work experience placement
    Local area
    Flexible hours

    Cisco

    San Jose, CA
    15 hours ago
  • $114.6k - $234.6k

     ...industry innovations to life-saving care. And with AI embedded across our products and services, we...  ...retrieval, storage, and processing. -Design performance and load testing.System Design & Architecture - System Reliability Design:-Build and design fault-tolerant components... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    10 hours ago
  • $146.3k - $306.4k

     ...architecture for large-scale systems software, firmware integration, and...  ...roadmaps to optimize for reliability, performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat...  ...to life-saving care. And with AI embedded across our products and... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    1 day ago
  • $167.7k - $245.2k

     ...TeamJoin Cisco's Enterprise AI team, the core group...  ...security — partnering across engineering, security, compliance,...  ...leader.As a Senior Software Engineer in Application Reliability, you will own the reliability...  ...to benchmark agent performance, test multi-step workflows... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • $155.8k - $224.2k

     ...a world powered by clean, reliable, and affordable energy is more...  ...revolutionizing power for AI-driven data centers to...  ...century.We are looking for a Principal Software Engineer, Middleware & Data to join...  ...Responsibilities:Develop high‑performance backend and systems components... 
    Performance
    Full time
    Work at office
    Worldwide

    Bloom Energy

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...manufacturers, and software providers to make inspection...  ....We’re seeking a Principal Systems Software Engineer for Semiconductor...  ..., multimodal AI, anomaly detection,...  ...memory, throughput, reliability, and security budgets...  ...about inference performance and evaluate hardware... 
    Performance
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $249k

     ...for travelers everywhere.Principal Software Development Engineer Our Technology Team partners...  ...without sacrificing reliability or skyrocketing our cloud...  ...: While we are embracing AI tools, your job is to build...  ...including capacity planning, performance optimization, and robust... 
    Performance
    Full time
    Weekend work

    Expedia

    San Jose, CA
    10 hours ago
  • $135.2k - $306.4k

     ...are hoping to enhance engineering efficiency by concentrating...  ...systems with high performance that can be adopted by...  ...that will ensure the reliability of databases being used...  ...As a Senior Principal Engineer, you will lead...  ...saving care. And with AI embedded across our products... 
    Performance
    Temporary work
    Worldwide
    Flexible hours

    Jobleads-US

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the next...  ...At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build...  ...the development of software systems. These...  ...needs. Establish reliability, security, validation...  ...silicon, or other high-performance computing systems.... 
    Performance
    Full time
    Remote work
    Shift work

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...seeking a highly motivated Principal System Software Engineer to drive next-generation...  ...system architecture, and performance engineering. In this highly...  ...hardware, architecture, kernel, AI, middleware, and platform...  ...to improve performance, reliability, determinism, and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $250.6k - $362.6k

     ...(CVIS) establishes and proves the performance of large AI clusters before they are handed over...  ...environments. Your Impact As a Principal Software Engineer, you will be the hands-on architect...  ...system capacity, delivery cycle time, reliability, or user adoption-and coordinating... 
    Performance
    Full time
    Temporary work
    Work experience placement
    Local area
    Flexible hours

    Cisco Systems, Inc.

    Milpitas, CA
    3 days ago
  • $272k - $431.25k

     ...environments. We are looking for Principal Software Engineers to help shape the...  ...operations, automation, and reliability across large-scale GPU clusters...  ...with GPU clusters, AI/ML infrastructure, Kubernetes...  ...Intelligence, High-Performance Computing and Visualization... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $104.5k - $234.6k

     ...architectures, efficient and reliable message brokering systems...  .... Career Level - IC4 Principal Platform Software Engineer. Lead platform projects...  ...upon completion. Conducts performance profiling and optimization...  ...life-saving care. And with AI embedded across our... 
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle

    Santa Clara, CA
    2 days ago
  • $207k - $300k

     ...members to enhance system reliability and efficiency....  ...reliability, scalability, and performance of SU services, often...  ...of experience with software development in one or...  ...as a Site Reliability Engineer.3 years of experience...  ...Experience in Generative AI, Generative AI Agent,... 
    Performance

    Google

    San Jose, CA
    10 hours ago
  • $272k - $431.25k

     ...strategic and technically proficient Principal Software Engineer to join the Data Center Systems and Software...  ...in designing scalable, high-performance server systems at the SW/HW interface...  ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA... 
    Performance
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    10 hours ago
  • $95k - $165k

     ...Segment: Home OfficeRole summary: As a Principal Software Engineer - Mobile, you will shape the next generation of AI-powered shopping experiences across Walmart...  ...evolving AI capabilities into intuitive, performant, and reliable customer experiences at Walmart scale.... 
    Performance
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    4 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  ...ROLE: AMD is looking for a Principal-level PyTorch training framework expert to help drive performance, scalability, and correctness...  ...communicate clearly with both engineers and stakeholders and can represent... 
    Performance

    AMD

    San Jose, CA
    10 hours ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline...  .... It combines software and systems engineering...  ...reliability and performance objectives while enabling...  ...environments.As a Principal SRE, you will shape the...  ...direction of NVIDIA’s AI Platform Runtime and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $132.6k - $214.5k

     ...Integrity, and Inclusion. We weave AI into the fabric of...  ...collaborate closely with our engineering teams to develop innovative...  ...insights into our systems’ performance and health. As a Senior Staff...  ...the product and ensure the reliability and availability of our services... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $156.4k - $253k

     ...Integrity, and Inclusion. We weave AI into the fabric of everything...  ...the Layer-7 Security Software team, we are responsible for...  ...Identification and Content Inspection Engine runs on Hardware, Virtualized...  ...terms of functionality and performance, working on device identity... 
    Performance
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer — AI Performance & Reliability. Be the first to apply!