Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer - Model Performance

Full-time

FATHOM

ABOUT FATHOM


We created Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments of your calls, so you and your team can stay fully present without sacrificing context or clarity. From instant, searchable call summaries to seamless CRM updates and team-wide sharing, Fathom transforms meetings from a source of friction into a place for alignment and momentum.

 

We’re a small company that creates magical experiences through the hard work of focused builders. We try to live our values - Care Deeply, Seek Leverage, Share Ownership, Sustain Urgency, and Be Tenacious - in everything we do, every day.

 

We started Fathom to rid us all of the tyranny of note-taking, and people seem to really love what we've built so far:

#1 Most Used App of the Year on HubSpot for 2025

Most installed AI meeting assistant on both the Zoom and HubSpot marketplaces

We’re hitting revenue and usage records every week

We think you’ll be pretty excited about Fathom too if you give it a try. Sign up today (it’s free)!

ROLE OVERVIEW

We're hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference stack, and to build the fine-tuning infrastructure that makes the rest of the AI team faster.

This is not a research role. You'll be optimizing real systems serving millions of meetings — choosing between quantization trade-offs, debugging speculative decoding, or figuring out why one GPU family's tail latency explodes at high concurrency while another stays stable.

You'll own two things:

1. Inference performance. You'll make our models faster and cheaper — speculative decoding, quantization, serving configuration, GPU selection, batching strategies, cold start mitigation, adapter swapping. Our traffic is extremely spiky (meetings end in 30-minute blocks), so you need to think about throughput curves. Our team greatly values offering a fast product.

2. Fine-tuning pipelines. The AI team constantly fine-tunes models for new tasks — distilling large teacher models for classification, training adapters for domain-specific behavior, DPO for preference tuning. Right now each project reinvents the training loop. You'll build repeatable infrastructure so an AI Engineer can go more quickly from dataset to deployed model.

HOW YOU’LL HELP US WIN

  • Benchmark FP8 quantization across GPU families, find that FP8 KV cache causes catastrophic repetition loops, identify static quantization as 6% faster than dynamic on certain hardware, and ship a production config that gets 1.3x speedup with <1% quality degradation

  • Evaluate serving frameworks (vLLM vs SGLang) with speculative decoding — discover that ngram speculation degrades ASR quality while EAGLE3 draft models don't, and that torch.compile makes certain GPUs 7% slower

  • Build a fine-tuning pipeline that takes a JSONL dataset and produces an optimized tune ready for serving, so a teammate can train a small classifier in an afternoon instead of a week

  • Optimize GPU spend — know which GPU families are best for batch workloads (stable under high concurrency) vs latency-sensitive paths (40% faster, but tail latency blows up under load), and when a 30% cost premium isn't worth it

  • Debug production inference issues — trace a quality regression to a serving framework upgrade that changed the default attention backend, or find that audio format handling in the multimodal pipeline silently drops segments

REQUIREMENTS

Hard Skills:

  • Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — not just deploying them, but tuning them: attention backends, scheduling strategies, CUDA graph warmup, prefix caching

  • Hands-on quantization experience — you've gone beyond "apply FP8 and hope." You understand weight vs activation quantization, per-channel vs per-tensor scaling, and when dynamic quantization introduces more overhead than it saves

  • Production fine-tuning experience — LoRA/QLoRA SFT, familiarity with training frameworks (ms-swift, Axolotl, torchtune, or similar), understanding of data formatting, learning rate schedules, and how to diagnose training failures

  • Strong Python. You'll write serving infrastructure, benchmarking harnesses, and training pipelines — not notebooks

  • Comfort with GPU profiling and performance analysis. You should be able to look at a benchmark result and know whether the bottleneck is compute, memory bandwidth, or scheduling overhead

Strong signal:

  • Cost modeling for GPU infrastructure — you've had to choose between GPU types and justify the tradeoff

  • Experience with multimodal models (audio/vision encoders + LLM decoders)

  • Experience with Modal, Ray Serve, or similar serverless GPU platforms

  • Understanding of audio processing (codecs, chunking, sample rates)

  • Experience building internal tooling that other engineers use — this role succeeds when the rest of the team ships faster

Not required:

  • ML research background or publications

  • Prompt engineering expertise (we have a team for that)

  • Frontend or full-stack experience

  • Masters/PhD (though it's fine if you have one)

 

WHAT'S IN IT FOR YOU

  • The opportunity to shape the foundational software services of a growing company

  • A role that balances innovation and incremental improvement

  • A dynamic and collaborative engineering team

  • Competitive compensation and benefits

  • A supportive environment that encourages innovation and personal growth

 

WHY YOU SHOULD JOIN US

  • Opportunity for impact. We’re established enough to ship instead of fighting fires and early enough that your work will have a real impact.

  • Startup experience. You’ll work closely with our CEO, a 2X Founder/CEO with a background in computer science and product design.

  • We embrace being fully remote. We schedule meetings sparingly and instead heavily use async comms (Slack, Notion, Loom)

ABOUT THE INTERVIEW

  • You’ll meet the entire team. We think it’s important that you get to meet everyone you’ll be working with.

  • No bullshit. Ask us anything you like. We’ve never understood why companies pretend they’re something that they’re not in the hiring process - you’re going to find out eventually so we’d rather you know who we are up front so we can both make sure this is a good fit for all involved.

  • Quick turnaround time. We know you have lots of options so we move fast usually in less than a week from start to finish.

HOW TO APPLY

Include a brief write-up or demo of inference optimization or model serving work you've done. We care about the reasoning behind your decisions — why you chose a specific quantization strategy, how you diagnosed a performance regression, what tradeoffs you navigated. A GitHub repo, blog post, or even a few paragraphs in your cover letter works.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Engineer - Model Performance in Remote vacancy
  • $125k - $150k

    Job DescriptionEverforth ECS is seeking an AI Model Engineer to work in a hybrid remote/onsite capacity, with minimum of 3 business days onsite...  ...deployment and monitoring processes while ensuring performance, observability, and security. This role contributes to building... 
    Performance
    Contract work
    Work at office
    Remote work

    ECS Federal

    Fairfax, VA
    1 day ago
  • $149 per hour

     ...provide more details.Vice President - Technical AI Foundation Model EngineerRole SummaryThe VP, Technical AI Foundation Model Engineer is responsible for designing, building,...  ...and recommend foundation models based on performance, cost, security, explainability, and... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    1 day per week

    MUFG

    Jersey City, NJ
    13 hours ago
  • $117.7k - $221.4k

     ...and cost efficient for embodied AI systems. We believe the next...  ...depends not only on stronger models, but also on better infrastructure...  ...-aware approach that first performs the cheapest reusable work, such...  ...model reflects how Cola engineers think: build durable intermediate... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    16 hours ago
  • $130k - $260k

     ...Opportunity: Walmart’s Supply Chain AI Lab & Innovation Factory is...  ...rigorous evaluation and model post-training. This role...  ...a deeply hands-on AI systems engineering position focused on designing...  ...and user trust.Set quality, performance, and release standards with layered... 
    Performance
    Full time
    Contract work
    Temporary work
    Part time

    Walmart

    Bentonville, AR
    16 hours ago
  •  ...now! Position Overview: The Staff AI/ML Engineer (LLMs) will lead the development of...  ...workflows Adapt and fine-tune foundation models for specialized use cases Design and...  ...program based on company and employee performance ~ Company paid life insurance, AD&D,... 
    Performance
    Full time
    Temporary work
    Work at office
    Visa sponsorship
    Relocation package
    Flexible hours

    Arka Group, L.p.

    Remote
    2 days ago
  •  ...leader in providing Information Technology, Engineering Services, Program Management, and...  ...Senior Machine Learning Engineers and AI Model Developers to support an upcoming Federal...  ...Model Selection Briefs documenting model performance, trade studies, reproducibility... 
    Performance
    For contractors
    Remote work
    Flexible hours

    Solerity

    Quantico, VA
    4 days ago
  • $165.2k - $223.6k

     ...unparalleled ML inference and training performance.The Inference Enablement and...  ...of running a wide range of models and supporting novel...  ...hardware-software boundary, our engineers build systematic...  ...boundaries of what's possible in AI acceleration.As part of the broader... 
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $114.6k - $252.1k

    Job Title: Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full timeMinimum Clearance Required to Start: TS/SCI with...  ...do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven by... 
    Contract work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CACI International

    Aurora, CO
    16 hours ago
  • $220k - $240k

    GovCIO is currently hiring an AI Ops Engineer with an active Secret clearance to implement artificial intelligence and automation capabilities...  ...and operational analytics.Improve system availability and performance through AI insights.QualificationsQualifications: Bachelor's... 
    Performance
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    16 hours ago
  •  ...Job Responsibilities: ML/AI fundamentals – Fine-tuning (SFT/preference), prompting, model-based eval, and the failure modes of each. At least one year of...  ...training/sampling jobs on real infra. Web/UI engineering – Production React/TypeScript, design, vibe code... 

    SGS Consulting

    Remote
    more than 2 months ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have...  ...closely with our AI/ML researchers to make their ML models more efficient leading to significant productivity improvements... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    16 hours ago
  • $80 - $95 per hour

     ...GorusuCompany: SRI Tech SolutionsJob Title : Senior AI Engineer No. Of Openings: 2Job Location: Houston,...  ...pipelines, emphasizing reliability, performance, and securityOrchestrate and configure...  ...AnsibleSolid understanding of MLOps, model lifecycle management, and CI/CD for AI... 
    Performance
    Hourly pay
    Full time
    Contract work
    For contractors
    Work from home

    SRI Tech

    Houston, TX
    4 days ago
  • $152.46k - $169.14k

     ...Qualifications Bachelor's degree in Engineering, or a related Science, Engineering or Mathematics...  ...information. Due to the nature of work performed within our facilities, U.S. citizenship...  ....Identifies opportunities to apply AI for continuous improvement and... 
    Performance
    Remote work
    Flexible hours

    General Dynamics Mission Systems

    Pittsfield, MA
    1 day ago
  • $122.57k - $204.25k

     ...Overview:We are looking for an AVP, IAM AI Engineer to design, operationalize, and automate...  ...grained, least-privilege authorization models for agents and workloads, expressed as...  ...salary of internal peers, demonstrated performance, and geographic location. Additionally,... 
    Performance
    Full time
    Work from home

    LPL Financial

    Austin, TX
    2 days ago
  • $141.7k - $268.3k

    In this position...The New Model Launch Manager holds a pivotal role in strategically overseeing and directing the...  ...success by leading, mentoring, and empowering a high-performing team of launch supervisors and engineers, fostering a high-performing organization. This... 
    Performance
    Immediate start
    Flexible hours

    Ford

    Allen Park, MI
    4 days ago
  • $152k - $241.5k

     ...deep learning ignited modern AI — the next era of computing —...  ...is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress...  ...end to end, with a focus on performance, fast builds, low memory use,...  ...hardware teams to enable new model patterns and upcoming GPU architectural... 
    Performance
    Full time
    Remote work

    Nvidia

    Austin, TX
    3 days ago
  • $120k - $220k

     ...information powered by advanced AI, recommendation systems, and...  ...every cycle.We're hiring the engineer who owns this agent end-to-end...  ...loop — Extend our LLM critic + performance-feedback regeneration from images...  ...0.50 per iteration.Generative model router — Pick the right model... 
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    16 hours ago
  • $112k - $179k

    ResponsibilitiesOverviewPeraton is seeking a Senior AI Engineer to to design and build production-grade...  ...real impact comes from orchestrating models, data, and workflows into production-...  ...closed-loop automationDefine and track performance metrics (cycle time, defect reduction,... 
    Performance
    Contract work
    Remote work
    Shift work

    Peraton Corporation

    Reston, VA
    2 days ago
  •  ...are seeking a Senior Forward Deployed AI Engineer to support our Public Sector initiatives...  ...responsible for transforming prototype models into scalable, efficient, and reliable...  ...Read hardware schematics/logs to identify performance bottlenecks and suggest improvements.... 
    Performance
    Full time
    Casual work
    Live out
    Work at office
    Local area
    Remote work

    webAI

    Austin, TX
    16 hours ago
  •  ...leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are...  ...is a team of researchers, engineers, designers, and more, who are all...  ...reusable tools for digging into model performance.You may be a good fit if:You... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    16 hours ago
  • $152k - $241.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ....We are looking for outstanding High Performance AI Engineer to build groundbreaking multi-agent systems...  ...workloads powered by foundational models. As a member of the team, you will develop... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Posting DescriptionPosition SummaryThe AI Engineer designs, develops, deploys, and supports...  ..., this role applies large language models (LLMs), agentic AI, machine learning (ML...  ...Implement AI governance processes, including performance monitoring, bias detection, drift... 
    Performance
    Full time
    Temporary work
    Remote work

    Boston Childrens’ Hospital

    Boston, MA
    3 days ago
  •  ..., apply now.We are currently seeking a AI Engineer to join our team in Edison, New Jersey...  ...development, and integration of complex AI models.• Work closely with data scientists,...  ...technologies into existing workflows.• Conduct performance evaluations of AI systems and provide... 
    Performance
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Edison, NJ
    1 day ago
  • OverviewThe AI Engineer is responsible for designing, developing, and deploying AI-driven...  ...be optimized. Leveraging large language models (LLMs) and agentic frameworks, the AI Engineer...  ...for efficient, effective, high-quality performance in self and in the department; delivers... 
    Performance
    Full time
    Local area
    Flexible hours

    PAM Health

    Plano, TX
    3 days ago
  • $99k - $225k

    AI EngineerThe Opportunity:As an AI engineer, you’ll design and deliver production-grade AI systems. You’ll own RAG...  ..., evaluation systems, and performance optimization leveraging modern AI...  ...prompt templates, prompt versioning, model configurations, and structured outputs... 
    Performance
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    McLean, VA
    3 days ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...areas, e.g. large language models, generative AI,...  ...must deliver leading inference performance, fast build time, reduced memory... 
    Performance
    Full time
    Remote work

    Nvidia

    Austin, TX
    16 hours ago
  • $99k - $225k

    Reinforcement Learning AI EngineerThe Opportunity:Are you an innovative...  ..., and machine learning (ML) engineering to train, test, deploy, and maintain models that learn from data to drive real...  ...Contribute to system architecture and performance optimization in Python with... 
    Performance
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    El Segundo, CA
    4 days ago
  • $99k - $225k

    AI EngineerThe Opportunity:As an AI Engineer, you will integrate AI-enabled capabilities into existing software systems by connecting models, inference endpoints, agent orchestration components, and...  ...environments. Due to the nature of work performed within this facility, U.S.... 
    Performance
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Huntsville, AL
    4 days ago
  • $225k - $250k

     ...CitiCitigroup Global Markets Inc. seeks a Model/Anlys/Valid Officer for its New York, New...  .... Improve the calculation speed and performance for production batch run to make sure trading...  ...individuals. Our automated processing and AI do not involve relying on automatic or... 
    Performance
    Full time
    Remote work

    Citigroup

    New York, NY
    16 hours ago
  • $86.8k - $198k

    AI Software Engineer, SeniorThe Opportunity:As an AI software engineer, you know that good software is more than just a nice-looking interface...  ...awards program acknowledges employees for exceptional performance and superior demonstration of our values. Full-time and part... 
    Performance
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Laurel, MD
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer - Model Performance. Be the first to apply!