Principal Software Engineer — AI Performance & Reliability
Jobleads-US
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.
Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.
THE ROLE:
We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.
Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.
This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.
You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.
THE PERSON:
- Profile and optimize AI model training and inference workloads.
- Improve model throughput, latency, memory efficiency, scalability, and reliability.
- Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
- Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
- Develop performance tooling, benchmarks, automation, and observability systems.
- Investigate and resolve complex production issues affecting AI workloads.
- Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.
- Translate customer feedback into product and infrastructure improvements.
- Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.
- Document performance findings, technical recommendations, and best practices.
KEY RESPONSIBILITIES:
- Strong software engineering skills and experience building production-quality systems.
- Experience working with AI infrastructure for model training, inference, or both.
- Demonstrated experience profiling and optimizing machine learning models or AI workloads.
- Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.
- Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication.
- Proficiency in languages such as Python, C++, or similar systems-oriented programming languages.
- Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
- Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack.
- Clear written and verbal communication skills.
- A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.
PREFERRED EXPERIENCE:
- Experience optimizing large language models, diffusion models, or recommendation models.
- Experience with GPU, accelerator, or distributed computing environments.
- Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.
- Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.
- Experience operating AI systems in production environments.
- Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role.
- Experience designing benchmarks and conducting systematic performance analysis.
ACADEMIC CREDENTIALS:
- A PhD (or a master’s degree with equivalent experience) in artificial intelligence, machine learning, computer science, or a related field.
LOCATION:
San Jose, CA or Bellevue, WA preferred (Hybrid). Other US locations may be considered.
#LI-MV1
#HYBRID
- Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.
This posting is for an existing vacancy.
#J-18808-Ljbffr Jobleads-US$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team... ...focal point for fleet-scale reliability, working directly with... ...Artificial Intelligence, High-Performance Computing and... ...existing vacancy. NVIDIA uses AI tools in its recruiting processes...PerformanceFull time$272k - $431.25k
...delivery across system software, drivers, and CUDA to... ...continuously available and reliable.What you’ll be doing:... .../platform layers, and performance counter/trace... ...integrate with existing ML/AI workflows (e.g., PyTorch... ...direction for an engineering team; mentor engineers...PerformanceFull time$272k - $431.25k
...unlimited potential of AI to define the next era... ...firmware and software architecture and design... ...passion for building reliable, debuggable, and scalable... ...Mentor architects and engineering teams to grow them into... ...architecture for scalable and performant edge systems,...PerformanceFull timeShift work$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as... ...to ensure they can reliably manage, update, and operate... ...secure boot, attestation), and performance — and champion those priorities... ...vacancy. NVIDIA uses AI tools in its recruiting processes...PerformanceFull timeRemote work$248k - $391k
...scalable simulation, AI, and thorough... ...systems that help engineers develop, test,... ...scenarios. In this Principal-level individual... ...parameter adaptation, software interfaces,... ...safety, simulation reliability, synthetic data,... ...dashboards, or high-performance numerical...PerformanceFull time$272k - $431.25k
...learning ignited modern AI—the next era of... ...assistants and engineering-productivity... ...company. Now we need a principal-level, hands-on... ...obsesses over reliability, polish, and user... ...Improve reliability, performance, observability,... ...like mature software, not prototypes.Build...PerformanceFull timeLive in$272k - $431.25k
...unlimited potential of AI to define the next era... ...maps improve driving performance, safety, and coverage.... ...with a diverse team of engineers in mapping, perception... ...fleet data into reliable map products used in self... ...building production-quality software systems.Solid...PerformanceFull timeWorldwide$221.2k - $387.1k
...DescriptionIt all started when engineer Fred Luddy wrote code... ..., ServiceNow is the AI control tower for... ...the quality and reliability of our security and risk... ...scalability, reliability, performance, and maintainability.... ...technical foundation in software architecture,...PerformanceWork at officeImmediate startRemote workFlexible hours$220k - $250k
...United StatesProducts - Engineering /Fulltime /HybridOver... ...of position: Principal Software EngineerPosition type... ...networking, generative AI, and autonomous agentic... ...be scalable, secure, reliable, observable, adaptable... ...networks by building high-performance, real-time systems...PerformanceFull timeH1bLocal areaWork from homeWork visaShift work$249k - $348.5k
...Principal Software Development Engineer Our Technology Team partners with teams across... ...by default. Hardened Reliability & Observability: Set SRE... ...direction. Familiarity with AI‑driven systems and... ...including capacity planning, performance optimization, and robust...PerformanceFlexible hours$250.6k - $362.6k
...Security team builds software and cloud product... ...team connects security engineering, security operations... ...Your Impact As a Principal Software Engineer,... ...distributed systems, reliability, and performance. You will mentor engineers... ...in the AI era – and beyond. We...PerformanceFull timeTemporary workWork experience placementLocal areaFlexible hours$114.6k - $234.6k
...industry innovations to life-saving care. And with AI embedded across our products and services, we... ...retrieval, storage, and processing. -Design performance and load testing.System Design & Architecture - System Reliability Design:-Build and design fault-tolerant components...PerformanceTemporary workFlexible hoursShift work$146.3k - $306.4k
...architecture for large-scale systems software, firmware integration, and... ...roadmaps to optimize for reliability, performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat... ...to life-saving care. And with AI embedded across our products and...PerformanceTemporary workFlexible hoursShift work$167.7k - $245.2k
...TeamJoin Cisco's Enterprise AI team, the core group... ...security — partnering across engineering, security, compliance,... ...leader.As a Senior Software Engineer in Application Reliability, you will own the reliability... ...to benchmark agent performance, test multi-step workflows...PerformanceFull timeTemporary workLocal areaFlexible hours$155.8k - $224.2k
...a world powered by clean, reliable, and affordable energy is more... ...revolutionizing power for AI-driven data centers to... ...century.We are looking for a Principal Software Engineer, Middleware & Data to join... ...Responsibilities:Develop high‑performance backend and systems components...PerformanceFull timeWork at officeWorldwide$272k - $431.25k
...manufacturers, and software providers to make inspection... ....We’re seeking a Principal Systems Software Engineer for Semiconductor... ..., multimodal AI, anomaly detection,... ...memory, throughput, reliability, and security budgets... ...about inference performance and evaluate hardware...PerformanceFull timeLocal areaShift work$249k
...for travelers everywhere.Principal Software Development Engineer Our Technology Team partners... ...without sacrificing reliability or skyrocketing our cloud... ...: While we are embracing AI tools, your job is to build... ...including capacity planning, performance optimization, and robust...PerformanceFull timeWeekend work$135.2k - $306.4k
...are hoping to enhance engineering efficiency by concentrating... ...systems with high performance that can be adopted by... ...that will ensure the reliability of databases being used... ...As a Senior Principal Engineer, you will lead... ...saving care. And with AI embedded across our products...PerformanceTemporary workWorldwideFlexible hours$272k - $431.25k
...unlimited potential of AI to define the next... ...At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build... ...the development of software systems. These... ...needs. Establish reliability, security, validation... ...silicon, or other high-performance computing systems....PerformanceFull timeRemote workShift work$272k - $431.25k
...seeking a highly motivated Principal System Software Engineer to drive next-generation... ...system architecture, and performance engineering. In this highly... ...hardware, architecture, kernel, AI, middleware, and platform... ...to improve performance, reliability, determinism, and...PerformanceFull time$250.6k - $362.6k
...(CVIS) establishes and proves the performance of large AI clusters before they are handed over... ...environments. Your Impact As a Principal Software Engineer, you will be the hands-on architect... ...system capacity, delivery cycle time, reliability, or user adoption-and coordinating...PerformanceFull timeTemporary workWork experience placementLocal areaFlexible hours$272k - $431.25k
...environments. We are looking for Principal Software Engineers to help shape the... ...operations, automation, and reliability across large-scale GPU clusters... ...with GPU clusters, AI/ML infrastructure, Kubernetes... ...Intelligence, High-Performance Computing and Visualization...PerformanceFull time$104.5k - $234.6k
...architectures, efficient and reliable message brokering systems... .... Career Level - IC4 Principal Platform Software Engineer. Lead platform projects... ...upon completion. Conducts performance profiling and optimization... ...life-saving care. And with AI embedded across our...PerformanceTemporary workFlexible hoursShift work$207k - $300k
...members to enhance system reliability and efficiency.... ...reliability, scalability, and performance of SU services, often... ...of experience with software development in one or... ...as a Site Reliability Engineer.3 years of experience... ...Experience in Generative AI, Generative AI Agent,...Performance$272k - $431.25k
...strategic and technically proficient Principal Software Engineer to join the Data Center Systems and Software... ...in designing scalable, high-performance server systems at the SW/HW interface... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA...PerformanceFull timeShift work$95k - $165k
...Segment: Home OfficeRole summary: As a Principal Software Engineer - Mobile, you will shape the next generation of AI-powered shopping experiences across Walmart... ...evolving AI capabilities into intuitive, performant, and reliable customer experiences at Walmart scale....PerformanceFull timeTemporary workPart time- ...generation computing experiences—from AI and data centers, to PCs,... ...ROLE: AMD is looking for a Principal-level PyTorch training framework expert to help drive performance, scalability, and correctness... ...communicate clearly with both engineers and stakeholders and can represent...Performance
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline... .... It combines software and systems engineering... ...reliability and performance objectives while enabling... ...environments.As a Principal SRE, you will shape the... ...direction of NVIDIA’s AI Platform Runtime and...PerformanceFull time$132.6k - $214.5k
...Integrity, and Inclusion. We weave AI into the fabric of... ...collaborate closely with our engineering teams to develop innovative... ...insights into our systems’ performance and health. As a Senior Staff... ...the product and ensure the reliability and availability of our services...PerformanceFull timeWork at officeVisa sponsorshipWork visa$156.4k - $253k
...Integrity, and Inclusion. We weave AI into the fabric of everything... ...the Layer-7 Security Software team, we are responsible for... ...Identification and Content Inspection Engine runs on Hardware, Virtualized... ...terms of functionality and performance, working on device identity...PerformanceFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer — AI Performance & Reliability. Be the first to apply!
- senior principal software engineer San Jose, CA
- principal software engineer San Jose, CA
- principal cloud computing engineer San Jose, CA
- senior principal cloud computing engineer San Jose, CA
- principal architect San Jose, CA
- principal San Jose, CA
- principal data scientist San Jose, CA
- senior principal scientist San Jose, CA
- ultimate software San Jose, CA
- software qa San Jose, CA


