Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Model Optimization Architect

$158.4k - $237.6k

Qualcomm

Company Qualcomm Technologies, Inc. Job Area Engineering Group, Engineering Group > Machine Learning Engineering General Summary Qualcomm is leveraging its strengths in compute, connectivity, and AI acceleration to play a central role in the evolution of Cloud AI. The Qualcomm Cloud AI team develops hardware and software platforms enabling efficient inference of large-scale foundation models. Position Overview We are seeking a Staff Engineer – AI Model Optimization Architect to lead end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators. This role works closely with compiler, performance, and accuracy teams to translate models into accelerator efficient execution while balancing throughput, latency, memory, and quality. The scope spans Day0 enablement through production deployment, with a strong emphasis on scaling optimizations to future architectures. Key Responsibilities Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators. Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations. Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites. Partner deeply with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration. Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes. Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance. Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency. Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems. Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling. Debug complex performance or stability issues to root cause and drive production ready solutions. Required Qualifications Expert level expertise in PyTorch and inference focused model optimization; strong Python engineering skills. Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows. Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs. Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs. Strong foundation in computer architecture, ML accelerators, and distributed systems. Proven ability to lead cross-functional technical efforts and influence design decisions. MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience. Preferred / Bonus Qualifications Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams. Familiarity with LLM serving stacks and continuous batching systems. Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks. PhD in a relevant field. Minimum Qualifications Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. Pay Range and Benefits $158,400.00 – $237,600.00. Salary is one component of total compensation. Competitive annual discretionary bonus program, opportunity for annual RSU grants. Highly competitive benefits package. For more details, refer to Qualcomm U.S. benefits information. Equal Opportunity Employer Qualcomm is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or any other protected classification. Qualcomm is committed to providing reasonable accommodations for individuals with disabilities. #J-18808-Ljbffr Qualcomm

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Model Optimization Architect in San Diego, CA vacancy
  • $158.4k - $237.6k

    Qualcomm is seeking a Staff Engineer - AI Model Optimization Architect in San Diego, California. This position involves leading end-to-end model transformation and optimization for large-scale models on Qualcomm accelerators, collaborating with compiler and performance... 
    Suggested

    Qualcomm

    San Diego, CA
    2 days ago
  • ServiceNow is seeking a GTM Business Process Optimization Manager to support the GTM Strategy and Process team through high-priority initiatives...  ...marketing, business development, and sales into one operating model, turning designs into actionable requirements for tech teams.... 
    Suggested

    ServiceNow

    San Diego, CA
    3 days ago
  • $105.78k - $189.35k

     ...want to be here!PURPOSE OF THE JOBThe AI Architect II is a senior technical leadership role...  ...covering RAG pipelines, agentic workflows, model serving, and AI gateway design.Lead the...  ...ensure consistency, security, and cost optimization.Collaborate with business and... 
    Suggested
    Full time
    Work at office
    Local area

    ICW Group

    San Diego, CA
    1 day ago
  •  ...a multi-modality foundation model to drive the next generation...  ...intelligence. As a Model Optimization & Deployment Engineer, you will...  ...-tuning (LoRA, QLoRA). Architect and implement model conversion...  ...maximize memory bandwidth on AI accelerators. Write production... 
    Suggested
    Temporary work
    Relocation package

    Zoox

    San Diego, CA
    23 days ago
  •  ...Specialist to own HRIS, benefits data, and related workflows. You'll optimize processes for direct-care staff and advance our UKG...  ...driven organization serving youth and families, with a hybrid work model in Oakland or San Diego and a focus on accurate data, strong #J... 
    Suggested
    Full time

    DeKalb Health

    San Diego, CA
    3 days ago
  •  ...and digital channels for inbound/outbound calls. The role includes leading the design of call flows, overseeing integrations, and optimizing system performance. The ideal candidate will have experience with telephony systems and cloud migration consulting, ensuring... 

    Kaleidoscope Innovation

    San Diego, CA
    2 days ago
  •  ...Description Title: Platform Architect/AWS solution Architect...  ...Practical experience with AWS AI/ML services (SageMaker, Bedrock...  ...Databricks Al capabilities (Model Serving, Feature Store, MLflow...  ..., re-architecting). • Cost Optimization: Ability to design cost-effective... 
    Shift work

    Krest Global Solutions

    San Diego, CA
    5 days ago
  • $180.4k - $270.6k

     ...(Multiple Levels) role at Qualcomm Get AI-powered advice on this job and more exclusive...  ...Work with low power team on power optimization Work with verification team to collaborate...  ...Specialist REMOTE Processor Performance Modeling Engineer Loan Partner - Client-Focused Mortgage... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home

    Qualcomm

    San Diego, CA
    2 days ago
  • $154.38k - $193.13k

    Job DescriptionAI Architect, AI & AutomationAbout the Role The applicant should have deep, hands-on experience architecting, designing, and...  ...multiple domains, including:Generative AI and Large Language Model (LLM) based solutionsPredictive and prescriptive Machine Learning... 
    Full time
    Temporary work
    Work experience placement

    Infosys Technologies

    San Diego, CA
    2 days ago
  • $172.8k - $304.9k

     ...Engineering Group, Engineering Group ASICS Engineering General Summary Qualcomm is looking for an experienced SoC architect to work on the next generation AI products in the datacenter. We are looking for a data center engineer whose expertise spans security, RAS and/... 
    Work experience placement
    Work from home

    Qualcomm

    San Diego, CA
    1 day ago
  •  ...pSemi Corporation, a Murata company, in San Diego, CA, seeks a Senior Staff Engineer to lead development of RFICs and AI-enabled platforms. You will drive intelligent automation for RF product development, define requirements, and coordinate cross-functional teams to... 

    Jobleads-US

    San Diego, CA
    1 day ago
  • $172.8k - $304.9k

     ...Qualcomm is seeking an experienced SoC architect to advance AI products in the datacenter. This role focuses on datacenter security, RAS, and virtualization, requiring significant expertise. Candidates must have extensive HW design experience and should excel in working... 

    Jobleads-US

    San Diego, CA
    3 days ago
  • Qualcomm Technologies, Inc. is seeking an experienced SoC Architect to define and deliver next-generation AI datacenter products and drive architecture across server SoC design, security, RAS, and virtualization. You will influence roadmaps and collaborate across hardware... 

    Qualcomm

    San Diego, CA
    4 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will create original, executable research problems that today’s frontier models cannot solve, sourced from papers, datasets, or your own scenarios... 
    Part time

    Mercor

    San Diego, CA
    2 days ago
  •  ...defines enterprise-grade governance frameworks, drives cross-functional alignment, and builds the operational infrastructure that keeps AI-powered content systems trustworthy, consistent, and scalable. You will set standards, build playbooks, partner with Engineering on... 

    ServiceNow

    San Diego, CA
    3 days ago
  •  ...with contractors and provisioning new hires from day one. The role is onsite and requires weekly inter-office commuting. You’ll drive AI-enabled automation to reduce repetitive work, implement robust identity and access controls, and design systems that scale as #J-188... 
    For contractors
    Work at office

    Triumph

    San Diego, CA
    3 days ago
  • Cadre AI is seeking a VP of AI Strategy to own the full lifecycle of AI strategy engagements, from kickoff through final delivery, and to convert completed projects into long-term retainers. You will engage directly with C-suite stakeholders, lead workshops, and produce... 

    Cadre AI

    San Diego, CA
    5 days ago
  • $153.47k - $255.75k

     ...seeking an experienced Identity and Access Management (IAM) Architect with a strong AI and agent-integration focus to lead the design, proof-of-...  ...Assess existing SSO, MFA, federation, and API authorization models; identify gaps in delegation, token lifecycle, scopes, secrets... 
    Work from home

    LPL Financial

    San Diego, CA
    1 day ago
  •  ...inform brand and marketing work, partnering across brand, product marketing, and research teams. You will translate complex market data into clear strategic recommendations for leadership, driving enterprise transformation and AI-driven branding. #J-18808-Ljbffr ServiceNow

    ServiceNow

    San Diego, CA
    2 days ago
  • Luminize is seeking a Business Process Automation Specialist in San Diego to own the automation layer for managing 80+ brands. Your role will involve designing, building, and maintaining custom tools and integrations to streamline operations. The ideal candidate has 5+ ...

    Luminize

    San Diego, CA
    21 hours ago
  •  ...Infosys Consulting seeks an AI Architect to design, develop, and deploy enterprise-grade AI/ML solutions across industries. You will lead cross-functional teams, translate business problems into AI architectures, and guide end-to-end project delivery with a focus on governance... 

    Jobleads-US

    San Diego, CA
    5 days ago
  •  ...years preferred Role Overview We are seeking a senior Enterprise AI Architect / AI Platform Architect to lead the design of enterprise AI...  ...reusable accelerators, platform standards, governance models, and enterprise architecture patterns. Leadership & Stakeholder... 

    Acunor

    San Diego, CA
    5 days ago
  • $126k - $229.8k

     ...Technologies, Inc. Engineering Group > ASICS Engineering General Summary Qualcomm is seeking an experienced SoC Architect to help define and deliver next-generation AI datacenter products. In this role, you will drive architecture and feature definition in key server... 
    Work experience placement
    Work from home

    Qualcomm

    San Diego, CA
    4 days ago
  • LPL Financial is seeking an experienced Identity and Access Management (IAM) Architect to lead the design and implementation of identity solutions for AI workflows. This role combines deep IAM expertise with engineering skills to produce secure identity standards in collaboration... 

    LPL Financial

    San Diego, CA
    21 hours ago
  • $105.78k - $189.35k

    A leading insurance firm in San Diego is looking for an AI Architect II to modernize its technology platform and lead AI initiatives. This senior technical role influences cloud and AI strategies while executing solutions. Responsibilities include establishing architectural... 

    ICW Group

    San Diego, CA
    5 days ago
  • $122.5k - $183.7k

    Qualcomm is seeking a Software Engineer specializing in computer vision systems. This role involves leading architecture decisions for advanced Snapdragon technology applications across multiple domains including mobile phones, autonomous vehicles, and robotics. The ideal...

    Qualcomm

    San Diego, CA
    21 hours ago
  • Job Overview:At Arm an SoC Architect is a technical role responsible for architecting and...  ...core, technology and software teams to optimize the end-to-end platform solutionsParticipate...  ...2.5D/3D packaging, performance / power modeling & estimation, soft real-time... 
    Work at office
    Local area

    ARM

    San Diego, CA
    4 days ago
  •  ...healthcare and scientific discovery to powering AI and the technologies people rely on every...  ...with passion to work on leading edge optimizing compilers for AMD GPU, we would love to...  ...understanding of GPU execution model and architectureACADEMIC CREDENTIALS:Bachelor... 

    AMD

    San Diego, CA
    1 day ago
  • $65 - $70 per hour

     ...WiproContact: Meghana GorusuCompany: SRI Tech SolutionsRole: Test Architect : (Priority 1)Location:San Diego, CA or Acton, MAOpening -1...  ...LoadRunner, or cloud-based tools; analyze bottlenecks and optimize application performanceCI/CD Integration: Embed test automation... 
    Hourly pay
    Full time

    SRI Tech

    San Diego, CA
    2 days ago
  • $90k - $125k

     ...seeking a motivated, collaborative Model-Based Systems Engineer to...  ...Modeler, Sparx Systems Enterprise Architect, IBM Rational Rhapsody, System...  ...and work-life balance. AI at G2 Ops. At G2 Ops, we don't...  ...into complex infrastructures, optimizing resiliency in system design and... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours

    G2 Ops Inc

    San Diego, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Model Optimization Architect. Be the first to apply!