AI Engineer - Model Performance
FATHOM
ABOUT FATHOM
We created Fathom to eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and organizes the key moments of your calls, so you and your team can stay fully present without sacrificing context or clarity. From instant, searchable call summaries to seamless CRM updates and team-wide sharing, Fathom transforms meetings from a source of friction into a place for alignment and momentum.
We’re a small company that creates magical experiences through the hard work of focused builders. We try to live our values - Care Deeply, Seek Leverage, Share Ownership, Sustain Urgency, and Be Tenacious - in everything we do, every day.
We started Fathom to rid us all of the tyranny of note-taking, and people seem to really love what we've built so far:
#1 Most Used App of the Year on HubSpot for 2025Most installed AI meeting assistant on both the Zoom and HubSpot marketplaces
We’re hitting revenue and usage records every week
We think you’ll be pretty excited about Fathom too if you give it a try. Sign up today (it’s free)!
ROLE OVERVIEWWe're hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference stack, and to build the fine-tuning infrastructure that makes the rest of the AI team faster.
This is not a research role. You'll be optimizing real systems serving millions of meetings — choosing between quantization trade-offs, debugging speculative decoding, or figuring out why one GPU family's tail latency explodes at high concurrency while another stays stable.
You'll own two things:
1. Inference performance. You'll make our models faster and cheaper — speculative decoding, quantization, serving configuration, GPU selection, batching strategies, cold start mitigation, adapter swapping. Our traffic is extremely spiky (meetings end in 30-minute blocks), so you need to think about throughput curves. Our team greatly values offering a fast product.
2. Fine-tuning pipelines. The AI team constantly fine-tunes models for new tasks — distilling large teacher models for classification, training adapters for domain-specific behavior, DPO for preference tuning. Right now each project reinvents the training loop. You'll build repeatable infrastructure so an AI Engineer can go more quickly from dataset to deployed model.
HOW YOU’LL HELP US WIN
Benchmark FP8 quantization across GPU families, find that FP8 KV cache causes catastrophic repetition loops, identify static quantization as 6% faster than dynamic on certain hardware, and ship a production config that gets 1.3x speedup with <1% quality degradation
Evaluate serving frameworks (vLLM vs SGLang) with speculative decoding — discover that ngram speculation degrades ASR quality while EAGLE3 draft models don't, and that torch.compile makes certain GPUs 7% slower
Build a fine-tuning pipeline that takes a JSONL dataset and produces an optimized tune ready for serving, so a teammate can train a small classifier in an afternoon instead of a week
Optimize GPU spend — know which GPU families are best for batch workloads (stable under high concurrency) vs latency-sensitive paths (40% faster, but tail latency blows up under load), and when a 30% cost premium isn't worth it
Debug production inference issues — trace a quality regression to a serving framework upgrade that changed the default attention backend, or find that audio format handling in the multimodal pipeline silently drops segments
REQUIREMENTS
Hard Skills:
Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — not just deploying them, but tuning them: attention backends, scheduling strategies, CUDA graph warmup, prefix caching
Hands-on quantization experience — you've gone beyond "apply FP8 and hope." You understand weight vs activation quantization, per-channel vs per-tensor scaling, and when dynamic quantization introduces more overhead than it saves
Production fine-tuning experience — LoRA/QLoRA SFT, familiarity with training frameworks (ms-swift, Axolotl, torchtune, or similar), understanding of data formatting, learning rate schedules, and how to diagnose training failures
Strong Python. You'll write serving infrastructure, benchmarking harnesses, and training pipelines — not notebooks
Comfort with GPU profiling and performance analysis. You should be able to look at a benchmark result and know whether the bottleneck is compute, memory bandwidth, or scheduling overhead
Strong signal:
Cost modeling for GPU infrastructure — you've had to choose between GPU types and justify the tradeoff
Experience with multimodal models (audio/vision encoders + LLM decoders)
Experience with Modal, Ray Serve, or similar serverless GPU platforms
Understanding of audio processing (codecs, chunking, sample rates)
Experience building internal tooling that other engineers use — this role succeeds when the rest of the team ships faster
Not required:
ML research background or publications
Prompt engineering expertise (we have a team for that)
Frontend or full-stack experience
Masters/PhD (though it's fine if you have one)
WHAT'S IN IT FOR YOU
The opportunity to shape the foundational software services of a growing company
A role that balances innovation and incremental improvement
A dynamic and collaborative engineering team
Competitive compensation and benefits
A supportive environment that encourages innovation and personal growth
WHY YOU SHOULD JOIN US
Opportunity for impact. We’re established enough to ship instead of fighting fires and early enough that your work will have a real impact.
Startup experience. You’ll work closely with our CEO, a 2X Founder/CEO with a background in computer science and product design.
We embrace being fully remote. We schedule meetings sparingly and instead heavily use async comms (Slack, Notion, Loom)
ABOUT THE INTERVIEW
You’ll meet the entire team. We think it’s important that you get to meet everyone you’ll be working with.
No bullshit. Ask us anything you like. We’ve never understood why companies pretend they’re something that they’re not in the hiring process - you’re going to find out eventually so we’d rather you know who we are up front so we can both make sure this is a good fit for all involved.
Quick turnaround time. We know you have lots of options so we move fast usually in less than a week from start to finish.
HOW TO APPLY
Include a brief write-up or demo of inference optimization or model serving work you've done. We care about the reasoning behind your decisions — why you chose a specific quantization strategy, how you diagnosed a performance regression, what tradeoffs you navigated. A GitHub repo, blog post, or even a few paragraphs in your cover letter works.
$125k - $150k
Job DescriptionEverforth ECS is seeking an AI Model Engineer to work in a hybrid remote/onsite capacity, with minimum of 3 business days onsite... ...deployment and monitoring processes while ensuring performance, observability, and security. This role contributes to building...PerformanceContract workWork at officeRemote work$149 per hour
...provide more details.Vice President - Technical AI Foundation Model EngineerRole SummaryThe VP, Technical AI Foundation Model Engineer is responsible for designing, building,... ...and recommend foundation models based on performance, cost, security, explainability, and...PerformanceFull timeWork at officeLocal areaRemote work1 day per week$117.7k - $221.4k
...and cost efficient for embodied AI systems. We believe the next... ...depends not only on stronger models, but also on better infrastructure... ...-aware approach that first performs the cheapest reusable work, such... ...model reflects how Cola engineers think: build durable intermediate...PerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$130k - $260k
...Opportunity: Walmart’s Supply Chain AI Lab & Innovation Factory is... ...rigorous evaluation and model post-training. This role... ...a deeply hands-on AI systems engineering position focused on designing... ...and user trust.Set quality, performance, and release standards with layered...PerformanceFull timeContract workTemporary workPart time- ...now! Position Overview: The Staff AI/ML Engineer (LLMs) will lead the development of... ...workflows Adapt and fine-tune foundation models for specialized use cases Design and... ...program based on company and employee performance ~ Company paid life insurance, AD&D,...PerformanceFull timeTemporary workWork at officeVisa sponsorshipRelocation packageFlexible hours
- ...leader in providing Information Technology, Engineering Services, Program Management, and... ...Senior Machine Learning Engineers and AI Model Developers to support an upcoming Federal... ...Model Selection Briefs documenting model performance, trade studies, reproducibility...PerformanceFor contractorsRemote workFlexible hours
$165.2k - $223.6k
...unparalleled ML inference and training performance.The Inference Enablement and... ...of running a wide range of models and supporting novel... ...hardware-software boundary, our engineers build systematic... ...boundaries of what's possible in AI acceleration.As part of the broader...PerformanceWork experience placementInternshipLocal areaFlexible hours$114.6k - $252.1k
Job Title: Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full timeMinimum Clearance Required to Start: TS/SCI with... ...do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven by...Contract workWork experience placementLocal areaRemote workFlexible hours$220k - $240k
GovCIO is currently hiring an AI Ops Engineer with an active Secret clearance to implement artificial intelligence and automation capabilities... ...and operational analytics.Improve system availability and performance through AI insights.QualificationsQualifications: Bachelor's...PerformanceCurrently hiringRemote work- ...Job Responsibilities: ML/AI fundamentals – Fine-tuning (SFT/preference), prompting, model-based eval, and the failure modes of each. At least one year of... ...training/sampling jobs on real infra. Web/UI engineering – Production React/TypeScript, design, vibe code...
$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have... ...closely with our AI/ML researchers to make their ML models more efficient leading to significant productivity improvements...PerformanceFull timeRemote work$80 - $95 per hour
...GorusuCompany: SRI Tech SolutionsJob Title : Senior AI Engineer No. Of Openings: 2Job Location: Houston,... ...pipelines, emphasizing reliability, performance, and securityOrchestrate and configure... ...AnsibleSolid understanding of MLOps, model lifecycle management, and CI/CD for AI...PerformanceHourly payFull timeContract workFor contractorsWork from home$152.46k - $169.14k
...Qualifications Bachelor's degree in Engineering, or a related Science, Engineering or Mathematics... ...information. Due to the nature of work performed within our facilities, U.S. citizenship... ....Identifies opportunities to apply AI for continuous improvement and...PerformanceRemote workFlexible hours$122.57k - $204.25k
...Overview:We are looking for an AVP, IAM AI Engineer to design, operationalize, and automate... ...grained, least-privilege authorization models for agents and workloads, expressed as... ...salary of internal peers, demonstrated performance, and geographic location. Additionally,...PerformanceFull timeWork from home$141.7k - $268.3k
In this position...The New Model Launch Manager holds a pivotal role in strategically overseeing and directing the... ...success by leading, mentoring, and empowering a high-performing team of launch supervisors and engineers, fostering a high-performing organization. This...PerformanceImmediate startFlexible hours$152k - $241.5k
...deep learning ignited modern AI — the next era of computing —... ...is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress... ...end to end, with a focus on performance, fast builds, low memory use,... ...hardware teams to enable new model patterns and upcoming GPU architectural...PerformanceFull timeRemote work$120k - $220k
...information powered by advanced AI, recommendation systems, and... ...every cycle.We're hiring the engineer who owns this agent end-to-end... ...loop — Extend our LLM critic + performance-feedback regeneration from images... ...0.50 per iteration.Generative model router — Pick the right model...PerformanceFull timeLocal areaWork from home$112k - $179k
ResponsibilitiesOverviewPeraton is seeking a Senior AI Engineer to to design and build production-grade... ...real impact comes from orchestrating models, data, and workflows into production-... ...closed-loop automationDefine and track performance metrics (cycle time, defect reduction,...PerformanceContract workRemote workShift work- ...are seeking a Senior Forward Deployed AI Engineer to support our Public Sector initiatives... ...responsible for transforming prototype models into scalable, efficient, and reliable... ...Read hardware schematics/logs to identify performance bottlenecks and suggest improvements....PerformanceFull timeCasual workLive outWork at officeLocal areaRemote work
- ...leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are... ...is a team of researchers, engineers, designers, and more, who are all... ...reusable tools for digging into model performance.You may be a good fit if:You...PerformanceFull timeWork at officeLocal areaRemote workHome office
$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ....We are looking for outstanding High Performance AI Engineer to build groundbreaking multi-agent systems... ...workloads powered by foundational models. As a member of the team, you will develop...PerformanceFull timeRemote work- ...Posting DescriptionPosition SummaryThe AI Engineer designs, develops, deploys, and supports... ..., this role applies large language models (LLMs), agentic AI, machine learning (ML... ...Implement AI governance processes, including performance monitoring, bias detection, drift...PerformanceFull timeTemporary workRemote work
- ..., apply now.We are currently seeking a AI Engineer to join our team in Edison, New Jersey... ...development, and integration of complex AI models.• Work closely with data scientists,... ...technologies into existing workflows.• Conduct performance evaluations of AI systems and provide...PerformanceWork at officeRemote workFlexible hours
- OverviewThe AI Engineer is responsible for designing, developing, and deploying AI-driven... ...be optimized. Leveraging large language models (LLMs) and agentic frameworks, the AI Engineer... ...for efficient, effective, high-quality performance in self and in the department; delivers...PerformanceFull timeLocal areaFlexible hours
$99k - $225k
AI EngineerThe Opportunity:As an AI engineer, you’ll design and deliver production-grade AI systems. You’ll own RAG... ..., evaluation systems, and performance optimization leveraging modern AI... ...prompt templates, prompt versioning, model configurations, and structured outputs...PerformanceFull timeContract workPart timeWork at officeLocal areaRemote work$152k - $241.5k
...into the unlimited potential of AI to define the next era of... ...AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers... ...areas, e.g. large language models, generative AI,... ...must deliver leading inference performance, fast build time, reduced memory...PerformanceFull timeRemote work$99k - $225k
Reinforcement Learning AI EngineerThe Opportunity:Are you an innovative... ..., and machine learning (ML) engineering to train, test, deploy, and maintain models that learn from data to drive real... ...Contribute to system architecture and performance optimization in Python with...PerformanceFull timeContract workPart timeWork at officeLocal areaRemote work$99k - $225k
AI EngineerThe Opportunity:As an AI Engineer, you will integrate AI-enabled capabilities into existing software systems by connecting models, inference endpoints, agent orchestration components, and... ...environments. Due to the nature of work performed within this facility, U.S....PerformanceFull timeContract workPart timeWork at officeLocal areaRemote work$225k - $250k
...CitiCitigroup Global Markets Inc. seeks a Model/Anlys/Valid Officer for its New York, New... .... Improve the calculation speed and performance for production batch run to make sure trading... ...individuals. Our automated processing and AI do not involve relying on automatic or...PerformanceFull timeRemote work$86.8k - $198k
AI Software Engineer, SeniorThe Opportunity:As an AI software engineer, you know that good software is more than just a nice-looking interface... ...awards program acknowledges employees for exceptional performance and superior demonstration of our values. Full-time and part...PerformanceFull timeContract workPart timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer - Model Performance. Be the first to apply!

