Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer - Model Performance

Full-time

FATHOM

ABOUT FATHOM


We created Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments of your calls, so you and your team can stay fully present without sacrificing context or clarity. From instant, searchable call summaries to seamless CRM updates and team-wide sharing, Fathom transforms meetings from a source of friction into a place for alignment and momentum.

 

We’re a small company that creates magical experiences through the hard work of focused builders. We try to live our values - Care Deeply, Seek Leverage, Share Ownership, Sustain Urgency, and Be Tenacious - in everything we do, every day.

 

We started Fathom to rid us all of the tyranny of note-taking, and people seem to really love what we've built so far:

#1 Most Used App of the Year on HubSpot for 2025

Most installed AI meeting assistant on both the Zoom and HubSpot marketplaces

We’re hitting revenue and usage records every week

We think you’ll be pretty excited about Fathom too if you give it a try. Sign up today (it’s free)!

ROLE OVERVIEW

We're hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference stack, and to build the fine-tuning infrastructure that makes the rest of the AI team faster.

This is not a research role. You'll be optimizing real systems serving millions of meetings — choosing between quantization trade-offs, debugging speculative decoding, or figuring out why one GPU family's tail latency explodes at high concurrency while another stays stable.

You'll own two things:

1. Inference performance. You'll make our models faster and cheaper — speculative decoding, quantization, serving configuration, GPU selection, batching strategies, cold start mitigation, adapter swapping. Our traffic is extremely spiky (meetings end in 30-minute blocks), so you need to think about throughput curves. Our team greatly values offering a fast product.

2. Fine-tuning pipelines. The AI team constantly fine-tunes models for new tasks — distilling large teacher models for classification, training adapters for domain-specific behavior, DPO for preference tuning. Right now each project reinvents the training loop. You'll build repeatable infrastructure so an AI Engineer can go more quickly from dataset to deployed model.

HOW YOU’LL HELP US WIN

  • Benchmark FP8 quantization across GPU families, find that FP8 KV cache causes catastrophic repetition loops, identify static quantization as 6% faster than dynamic on certain hardware, and ship a production config that gets 1.3x speedup with <1% quality degradation

  • Evaluate serving frameworks (vLLM vs SGLang) with speculative decoding — discover that ngram speculation degrades ASR quality while EAGLE3 draft models don't, and that torch.compile makes certain GPUs 7% slower

  • Build a fine-tuning pipeline that takes a JSONL dataset and produces an optimized tune ready for serving, so a teammate can train a small classifier in an afternoon instead of a week

  • Optimize GPU spend — know which GPU families are best for batch workloads (stable under high concurrency) vs latency-sensitive paths (40% faster, but tail latency blows up under load), and when a 30% cost premium isn't worth it

  • Debug production inference issues — trace a quality regression to a serving framework upgrade that changed the default attention backend, or find that audio format handling in the multimodal pipeline silently drops segments

REQUIREMENTS

Hard Skills:

  • Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — not just deploying them, but tuning them: attention backends, scheduling strategies, CUDA graph warmup, prefix caching

  • Hands-on quantization experience — you've gone beyond "apply FP8 and hope." You understand weight vs activation quantization, per-channel vs per-tensor scaling, and when dynamic quantization introduces more overhead than it saves

  • Production fine-tuning experience — LoRA/QLoRA SFT, familiarity with training frameworks (ms-swift, Axolotl, torchtune, or similar), understanding of data formatting, learning rate schedules, and how to diagnose training failures

  • Strong Python. You'll write serving infrastructure, benchmarking harnesses, and training pipelines — not notebooks

  • Comfort with GPU profiling and performance analysis. You should be able to look at a benchmark result and know whether the bottleneck is compute, memory bandwidth, or scheduling overhead

Strong signal:

  • Cost modeling for GPU infrastructure — you've had to choose between GPU types and justify the tradeoff

  • Experience with multimodal models (audio/vision encoders + LLM decoders)

  • Experience with Modal, Ray Serve, or similar serverless GPU platforms

  • Understanding of audio processing (codecs, chunking, sample rates)

  • Experience building internal tooling that other engineers use — this role succeeds when the rest of the team ships faster

Not required:

  • ML research background or publications

  • Prompt engineering expertise (we have a team for that)

  • Frontend or full-stack experience

  • Masters/PhD (though it's fine if you have one)

 

WHAT'S IN IT FOR YOU

  • The opportunity to shape the foundational software services of a growing company

  • A role that balances innovation and incremental improvement

  • A dynamic and collaborative engineering team

  • Competitive compensation and benefits

  • A supportive environment that encourages innovation and personal growth

 

WHY YOU SHOULD JOIN US

  • Opportunity for impact. We’re established enough to ship instead of fighting fires and early enough that your work will have a real impact.

  • Startup experience. You’ll work closely with our CEO, a 2X Founder/CEO with a background in computer science and product design.

  • We embrace being fully remote. We schedule meetings sparingly and instead heavily use async comms (Slack, Notion, Loom)

ABOUT THE INTERVIEW

  • You’ll meet the entire team. We think it’s important that you get to meet everyone you’ll be working with.

  • No bullshit. Ask us anything you like. We’ve never understood why companies pretend they’re something that they’re not in the hiring process - you’re going to find out eventually so we’d rather you know who we are up front so we can both make sure this is a good fit for all involved.

  • Quick turnaround time. We know you have lots of options so we move fast usually in less than a week from start to finish.

HOW TO APPLY

Include a brief write-up or demo of inference optimization or model serving work you've done. We care about the reasoning behind your decisions — why you chose a specific quantization strategy, how you diagnosed a performance regression, what tradeoffs you navigated. A GitHub repo, blog post, or even a few paragraphs in your cover letter works.

Vacancy posted 23 hours ago
Similar jobs that could be interesting for youBased on the AI Engineer - Model Performance in Remote vacancy
  •  ...thinking organization, apply now.We are currently seeking a AI Foundational Model Engineer to join our team in Jersey City, New Jersey (US-NJ),...  ...incentive compensation based on individual and/or company performance. If the position offered in temporary, the position will... 
    Performance
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Jersey City, NJ
    4 days ago
  • $117.7k - $221.4k

     ...and cost efficient for embodied AI systems. We believe the next...  ...depends not only on stronger models, but also on better infrastructure...  ...-aware approach that first performs the cheapest reusable work, such...  ...model reflects how Cola engineers think: build durable intermediate... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...now! Position Overview: The Staff AI/ML Engineer (LLMs) will lead the development of...  ...workflows Adapt and fine-tune foundation models for specialized use cases Design and...  ...program based on company and employee performance ~ Company paid life insurance, AD&D,... 
    Performance
    Full time
    Temporary work
    Work at office
    Visa sponsorship
    Relocation package
    Flexible hours

    Arka Group, L.p.

    Remote
    23 hours ago
  •  ...Position Overview: The Principal AI/ML Engineer will support the development of AI/ML algorithms...  ...learning, and large language models. We offer generous relocation benefits...  ...program based on company and employee performance ~ Company paid life insurance, AD&D,... 
    Performance
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Arka Group, Lp

    Remote
    23 hours ago
  • $125k - $150k

     ...DescriptionEverforth ECS is seeking an AI Model Developer to work in a hybrid remote/onsite...  ...role requires the need to build high performance Python-based AI/ML solutions capable of...  .... Perform data preprocessing, feature engineering, and model optimization. Evaluate... 
    Performance
    Contract work
    Work at office
    Remote work

    ECS Federal

    Fairfax, VA
    4 days ago
  •  ...leader in providing Information Technology, Engineering Services, Program Management, and...  ...Senior Machine Learning Engineers and AI Model Developers to support an upcoming Federal...  ...Model Selection Briefs documenting model performance, trade studies, reproducibility... 
    Performance
    For contractors
    Remote work
    Flexible hours

    Solerity

    Fort Belvoir, VA
    5 days ago
  • $129.6k - $244.68k

     ...have...Bachelor’s degree in mechanical engineering and at least 5 years of experience or equivalent...  ...Developing full vehicle impact models using software ANSA, PRIMER, HYPERWORKS...  ...LS Dyna Debugging models, root-cause performance issues, and identifying solutions or alternative... 
    Performance
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford

    Allen Park, MI
    2 days ago
  •  ...financial world.The roleSoFi is seeking a Fraud Model Developer to join our Fraud Model...  ...You will also analyze model and product performance, identify key drivers of fraud losses,...  ...Fraud Risk, Fraud Operations, Product, Engineering, Finance, Accounting, and other business... 
    Performance
    Remote work

    SoFi

    Frisco, TX
    5 days ago
  • $114.6k - $252.1k

    Job Title: Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full timeMinimum Clearance Required to Start: TS/SCIEmployee...  ...do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven... 
    Contract work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CACI International

    Aurora, CO
    23 hours ago
  • $165.2k - $223.6k

     ...unparalleled ML inference and training performance.The Inference Enablement and...  ...of running a wide range of models and supporting novel...  ...hardware-software boundary, our engineers build systematic...  ...boundaries of what's possible in AI acceleration.As part of the broader... 
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...company is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This...  ...and some coding knowledge. You will contribute to training AI models while receiving competitive hourly pay. Flexible schedules and project... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    3 days ago
  • $40 per hour

    A leading tech company is seeking experienced cybersecurity professionals for a remote role to help train AI models. You will evaluate AI-generated cybersecurity content, solve technical problems, and provide feedback to enhance AI systems. The ideal candidate has over... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    3 days ago
  • $40 per hour

     ...experienced cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content...  ...testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or similar) ~ Some... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Washington DC
    3 days ago
  • $40 per hour

    A data and AI solutions provider is seeking experienced cybersecurity professionals for remote roles to evaluate AI-generated security...  ...will work on projects that directly influence AI security model development. Opportunities available for candidates in multiple countries... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Virginia, MN
    1 day ago
  • $220k - $240k

    GovCIO is currently hiring an AI Ops Engineer with an active Secret clearance to implement artificial intelligence and automation capabilities...  ...and operational analytics.Improve system availability and performance through AI insights.QualificationsQualifications: Bachelor's... 
    Performance
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  • $40 per hour

    A technology company in the United States is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve cybersecurity problems. This can be a full-time or part-time remote position with hourly pay starting at $40+. Candidates should... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Indiana, PA
    3 days ago
  •  ...Support Internal Audit’s evaluation of model and artificial intelligence (AI) risk and governance frameworks...  ..., Mathematics, Computers Science, Engineering, or degrees in similar...  ...materiality of model changes, ongoing performance monitoring, and other targeted model... 
    Performance
    Internship
    Monday to Friday

    Navy Federal Credit Union

    Vienna, VA
    1 day ago
  • $40 per hour

    A tech company specializing in AI security is seeking experienced cybersecurity professionals for a remote role. Responsibilities include evaluating AI-generated security content and solving technical problems related to cybersecurity. Candidates should have 2+ years of... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    El Paso, TX
    3 days ago
  • $40 per hour

    A leading AI training company in the United States is seeking experienced cybersecurity professionals for a remote role. The job involves evaluating AI-generated security content and solving technical challenges to improve AI systems. Ideal candidates will have over 2... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Providence, RI
    3 days ago
  • $80 - $95 per hour

     ...GorusuCompany: SRI Tech SolutionsJob Title : Senior AI Engineer No. Of Openings: 2Job Location: Houston,...  ...pipelines, emphasizing reliability, performance, and securityOrchestrate and configure...  ...AnsibleSolid understanding of MLOps, model lifecycle management, and CI/CD for AI... 
    Performance
    Hourly pay
    Full time
    Contract work
    For contractors
    Work from home

    SRI Tech

    Houston, TX
    23 hours ago
  • DescriptionAI Engineer - Financial Services Remote / HybridAbout RiskSpanRiskSpan...  ...leading source of analytics, modeling, data, and risk management...  ...We are seeking a hands-on AI Engineer to design, build,...  ...guardrails.· Evaluate model performance and iteratively improve... 
    Performance
    Remote work

    RiskSpan

    Washington DC
    4 days ago
  • $40 per hour

    A leading AI cybersecurity firm is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, tackle technical cybersecurity challenges, and provide crucial feedback to enhance AI systems. The ideal candidate... 
    Hourly pay
    Remote work

    DataAnnotation

    Honolulu, HI
    4 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have...  ...closely with our AI/ML researchers to make their ML models more efficient leading to significant productivity improvements... 
    Performance
    Full time
    Remote work

    Nvidia

    Seattle, WA
    1 day ago
  •  ...FinTechSelling Points Drive innovation in AI systems in a hybrid work...  ...solutions.Mentor and guide engineering teams, fostering growth and...  ...and improve AI system performance.Enhance team processes and infrastructure...  ...MLOps practices, including model training, deployment, and... 
    Performance
    Work at office
    Remote work

    Green Key Resources

    New York, NY
    1 day ago
  • $152k - $241.5k

     ...GPUs are at the core of modern AI infrastructure, from training large-scale models to running inference in production...  ...much as hardware, and compiler engineering is a big part of what makes it work...  ...passes, and target-specific performance signals.Apply RL techniques to optimize... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $70 - $75 per hour

     ...PARTIES, or VISA SPONSORSHIPPosition Title: AI Engineer (Automation & Intelligent Systems)...  ...automation opportunities.Integrate AI/ML models into automation workflows to enable...  ...and scalable automation systems.Monitor performance, troubleshoot complex issues, and ensure... 
    Performance
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    VACO

    Dublin, OH
    5 days ago
  • $65.1k - $162.12k

     ...That fighting spirit lives on in Ford Racing today. We're the engineers, strategists, and competitors who bring Ford's track-tested edge...  ...win, engineer to lead.If you're a competitor, innovator, or performance obsessive, Ford Racing is where you belong. Help us write the... 
    Performance
    Immediate start

    Ford

    Allen Park, MI
    4 days ago
  • $40 per hour

    A cybersecurity solutions company is seeking experienced professionals to join a remote team. In this role, you'll evaluate AI-generated cybersecurity content and solve technical security problems. Candidates should have over 2 years of hands-on cybersecurity experience... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    3 days ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...areas, e.g. large language models, generative AI,...  ...must deliver leading inference performance, fast build time, reduced memory... 
    Performance
    Full time
    Remote work

    Nvidia

    Austin, TX
    1 day ago
  • $152k - $241.5k

    We’re currently seeking a Senior AI Developer Technology Engineer, Financial Sector!Would you like to help...  ...to achieve the best possible performance of computer hardware? Could you be thrilled...  ..., software, and programming models in collaboration with research, hardware... 
    Performance
    Full time
    Work experience placement
    Remote work

    Nvidia

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer - Model Performance. Be the first to apply!