Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Back End Engineer, Evals - Hazel AI

$275k - $325k

Altruist

About Altruist

Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time and resources, and deliver superior outcomes for their clients.

We're looking for exceptional talent to help us achieve our mission of making financial advice better, more affordable, and accessible to all. If you're passionate about challenging the status quo and want to do the most important work of your life, we'd love to meet you!

But first, our values

Kindness - Kindness doesn't just equal niceness. We listen to understand. We embrace, and encourage healthy debate and diverse perspectives. We approach conflict openly, honestly, and respectfully.

Brilliance - Humility is the skill we're most proud of and possessing a growth mindset is always top of mind. We take ownership in everything we touch; regularly using our unique superpowers to reach a common goal as a team. We succeed and fail as one.

Grit - When challenges arise, we stay laser focused on achieving our mission and finding a way forward, even when it's hard. We are nimble and maintain a sense of urgency, swiftly adapting to change and overcoming obstacles.

About Hazel:

Hazel.ai is building the AI engine for wealth management that unlocks 10x growth, efficiency and value for financial advisors and their clients in a regulated industry. Since its launch last September, Hazel has organically and rapidly grown its user base.

Hazel is a part of Altruist's broader mission to make financial advice better, more affordable, and accessible to all.

This role is hybrid, with four in-office days per week at our San Francisco FiDi location.

The opportunity:

Architect our evaluation platform from first principles - the observability, scoring, golden datasets, verification agents, and CI/CD integration that define standards of quality. You'll work shoulder-to-shoulder with backend engineers, product managers, and a growing bench of subject matter experts, including practicing CFPs, CPAs, and tax planners, to translate fiduciary-grade requirements into automated quality signals.

Your impact:
  • Design and build Hazel's evals platform end-to-end - online scoring, offline benchmarks, regression suites, LLM-as-judge pipelines, and human-in-the-loop review workflows across every Hazel surface.
  • Build production observability and monitoring for AI quality: hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows.
  • Architect data curation pipelines that turn real advisor interactions into evaluation datasets - with rigorous sampling strategies, labeling protocols, dataset versioning, and the privacy and consent controls required for regulated finance.
  • Build and steward Hazel's golden datasets in close partnership with SMEs and a network of practicing advisors, CFPs, and tax professionals - translating their tacit expertise into precise, measurable eval criteria.
  • Develop LLM verification agents that catch hallucinations, computational errors, and compliance violations before they ever reach an advisor or client.
  • Integrate evals into our deployment pipeline so that every prompt change, model swap, harness modification, or RAG pipeline tweak runs against regression and acceptance criteria before shipping - making evals a first-class deployment gate, not a quarterly audit.
  • Partner with the team building Hazel's model-agnostic orchestration harness to evaluate cross-model and cross-provider performance, surface tradeoffs, and inform routing decisions across Anthropic, OpenAI, and self-hosted models.
  • Define quality SLOs for each Hazel surface and build alerting that catches regressions in production before our customers do - especially for high-stakes flows like tax and financial planning.
  • Establish Hazel's eval methodology as a defensible competitive advantage - infrastructure good enough that model upgrades from frontier labs become accelerants for us, not threats.
What you bring:
  • 8+ years of engineering experience, with at least 2 years focused on evaluation infrastructure, model quality, fine-tuning, or ML platform work for production systems.
  • Deep familiarity with evaluation and scoring methodologies for modern AI systems - RAG evaluation, document processing, fine-tuned model assessment, agentic and tool-use system evaluation, LLM-as-judge frameworks, and human evaluation protocols.
  • Experience designing and curating golden datasets - sampling strategies, inter-rater agreement, dataset versioning, and managing the long tail of edge cases.
  • Comfort working across the stack - data engineering (SQL, dbt, warehouses), backend integration (APIs, async pipelines, queues), and observability tooling.
  • Strong communication skills. You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria - and explain quality tradeoffs clearly to engineers, product managers, and leadership.
  • A bias toward shipping. You believe great evals enable speed, not just safety, and you build tools that engineers actually want to use.
  • Bonus Points:
    • Prior experience at an applied AI company building evals, model quality, or applied research infrastructure.
    • Experience evaluating multi-step agentic workflows, tool-use systems, or RAG pipelines in production.
    • Familiarity with frameworks like Braintrust, Langfuse or similar - including a clear point of view on when to use which.
    • Background in regulated industries (financial services, healthcare, legal) where accuracy, auditability, and the cost of a wrong answer are unusually high.
    • Experience building human-in-the-loop labeling workflows, annotation tooling, or red-teaming programs.
    • Domain knowledge of wealth management, tax planning, or financial planning - or genuine excitement to learn it deeply alongside our SME bench.
San Francisco, CA salary range

$275,000-$325,000 USD

What we bring

Attracting and retaining top-tier talent is a priority. We are proud of the culture we've built and are cognizant of the ever-changing professional landscape. Our dynamic offering of perks and benefits are tailored for you to feel your best while doing your best.
  • Stunning, amenity-filled office spaces in Culver City, CA, San Francisco, CA, and Dallas, TX. Our offices are intentionally designed for comfort, collaboration, and productivity.
  • Competitive pay and equity for eligible positions.
  • Premium healthcare, dental, and vision insurance plans (HMO and PPO).
  • 401k savings plan with a 4% match and immediate vesting.
  • 16 week paid parental leave after one year of employment.
  • Professional growth and development opportunities including an employee mobility program and an annual L&D budget allocation for each employee.
  • Company perks program (includes discounts on pet insurance, fitness, cell phone plans, and travel, etc.).
  • Financial guidance program (includes counseling on navigating debt, tracking personal spend, saving and planning goals, home-purchasing preparedness, etc.).
  • One month work from anywhere policy (with the exception of a few countries).

Total compensation includes a competitive benefits package, along with equity in the form of Stock Options (ISOs) for eligible roles. For salaried positions, a salary offer will be determined by a number of factors including experience, skill level, internal pay equity, geographic location, and other relevant business considerations. We review all employee pay and compensation programs regularly to ensure fair, equitable, and competitive pay. At Altruist, we are committed to providing fair, equitable, and competitive compensation by leveraging market data to inform our pay bands. Base salaries will be reviewed at regular intervals throughout the year, typically in conjunction with performance review cycles. By evaluating compensation on a regular basis, we are able to reward high performance and ensure all employees have opportunities for growth.

Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. At Altruist we are dedicated to building a diverse, inclusive, and authentic workplace, so if you're excited about this role, but your past experience doesn't align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.
Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Staff Back End Engineer, Evals - Hazel AI in San Francisco, CA vacancy
  • $230k - $385k

     ...organization by applying cutting-edge AI models to real-world...  .... From customer operations to engineering, we develop an ecosystem of automation...  ...help to design and build an evals infrastructure that measures...  ...termination of employment or end of assignment; and maintain... 
    Suggested
    Internship
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $140k - $175k

     ...developers build mission-critical AI applications across the entire...  ...in 2023, LangChain powers top engineering teams at companies like Replit...  ...commercial observability and evals platform product. In this role...  ...customers, developer end-users and internal stakeholders... 
    Suggested
    Full time

    LangChain

    San Francisco, CA
    20 hours ago
  • $145k - $180k

     ...build the foundation for agent engineering in the real world, helping...  ...prototypes to production-ready AI agents that teams can rely on....  ...commercial AI observability and evals platform product. In this role...  ...enterprise customers, developer end-users and internal... 
    Suggested
    Full time
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    20 hours ago
  • $175k - $240k

     ...build the foundation for agent engineering in the real world, helping...  ...prototypes to production-ready AI agents that teams can rely on....  ...LangSmith, an observability and evals platform. In this role, you'll...  ...enterprise customers, developer end-users and internal... 
    Suggested
    Full time
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    20 hours ago
  •  ...Michael Antonov, co-founder of Oculus, and backed by Formic Ventures, we are redefining...  ...infrastructure behind modern drug discovery. Our AI-driven platform enables scientists to...  ...are seeking a talented and experienced Staff Engineer to join our dynamic team. You will be a... 
    Suggested
    Temporary work
    Work at office
    Work visa
    3 days per week

    Deep Origin

    South San Francisco, CA
    3 days ago
  • $140k - $260k

     ...help companies understand and control their AI presence. We're building the foundational...  ...What you’ll do Build core workflow engine primitives used to orchestrate agents, tools...  ...as ideal, and comfortable owning services end to end in production Solid experience with... 
    Full time
    Work at office
    Visa sponsorship

    Profound

    San Francisco, CA
    20 hours ago
  •  ...Lantern is building an AI radiology resident-think...  ...As a backend software engineer working on our radiology...  ...you'll do Take on end-to-end ownership for delivering...  ...radiologists, clinical staff, and IT staff to...  ...products or systems, including evals, fine tuning, operations... 
    Full time

    New Lantern

    San Francisco, CA
    20 hours ago
  •  ...Monitoring Platform Raindrop is the monitoring platform for AI agents. Engineering teams at Fortune 100s and the fastest-growing AI companies...  ...sifting through millions of logs and trying to debug flaky evals that just aren't matching real world results. Evals are like... 
    Temporary work

    Raindrop

    San Francisco, CA
    1 day ago
  • $139k - $257.55k

     ...creativity — where generative AI, intelligent agents, and human...  ...hand in hand. As a Full Stack Engineer, Agentic Product, you will be...  ...framework for agent behavior — the evals, benchmarks, and systematic...  ...realities of running LLM-backed systems in production. Experience... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    1 day ago
  •  ...Backend Engineer Arena Intelligence is looking for a backend engineer to build the products...  ..., and data systems that turn Arena's evals into products that enterprises depend on....  ...and Orb. Familiarity with the modern AI stack (vLLM, LiteLLM, LangChain, etc.).... 
    Permanent employment
    Flexible hours
    Shift work

    Arena AI

    San Francisco, CA
    20 hours ago
  •  ...About Reducto Reducto helps AI teams ingest real world enterprise data with state of...  ...Capital, and are looking for a Backend/AI Engineer. The Opportunity As a Backend/AI Engineer...  ...latency Building internal tooling and evals to better understand/analyze failure cases... 
    Full time
    Work at office
    Local area

    Reducto

    San Francisco, CA
    20 hours ago
  • $100k - $130k

     ...SaaS redefining responsible data use for the AI era. The Ketch Data Permissioning Platform...  ...empowers businesses and software engineers with a deploy-once, comply-and-secure-everywhere...  ...customers to understand use cases and solutions end-to-end Collectively be responsible for... 
    Full time

    Ketch

    San Francisco, CA
    1 hour ago
  • $120k - $250k

     ...Overview Founding Full-Stack Engineer role at Floot (YC S25). Base pay...  ...hats and touch everything from AI systems, user-facing features,...  ...reliable Prompt engineering and evals to improve AI output quality (for...  ...robust infrastructure for our end-to-end platform to power millions... 
    Full time

    Floot (YC S25)

    San Francisco, CA
    3 days ago
  • PLEASE CLICK HERE TO SEE *ALL* OF OUR JOB OPENINGS!Senior Full Stack Engineer (Backend)As a Senior Full-Stack Engineer, you'll be a key player...  ...insights, working with cutting-edge algorithms and composite AI solutions. This role offers the perfect balance of technical... 
    Casual work
    Work at office

    Three Pillars Recruiting

    San Francisco, CA
    2 days ago
  •  ...OUR JOB OPENINGS!Senior Backend EngineerSeeking a Senior Backend Engineer to architect and operate their backend platform that delivers...  ...complex problems your wayCutting-edge technology: Work with advanced AI/ML algorithms, composite AI solutions, private NVIDIA DGX... 
    Casual work

    Three Pillars Recruiting

    San Francisco, CA
    2 days ago
  •  ...grade security isn't just available off the shelf. That's where the AI security team comes in—we're responsible for building and...  ...priority is a must.This is a foundational role—you'd be the first engineer on the team.What you'll doEngineers on this team do a mixture of... 

    Stripe

    San Francisco, CA
    1 day ago
  • $166.6k - $208.3k

     ...navigation through the skies. It embodies the elegance of simplicity in engineering, transforming the demanding task of controlling an aircraft...  ...technical opinions and lay out tradeoffs.Actively incorporates AI tools into their workflow and is excited about leveraging them... 

    Mercury

    San Francisco, CA
    16 hours ago
  • $140k - $190k

     ...siteLoft Orbital is looking for a Backend Engineer to join our Mission Operations Services (MOS...  ...backend code in Python, owning services end‑to‑end from design and implementation through...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more.... 
    Full time
    Local area

    Loft Orbital

    San Francisco, CA
    2 days ago
  •  ...source, vet, and onboard expert contractors who help train AI models in a wide variety of domains. Our technology is so...  ...people. About the Role As a Senior Software Engineer at Mercor, you’ll take end-to-end ownership of complex projects that power the workflows... 
    Full time
    For contractors
    Immediate start

    Mercor

    San Francisco, CA
    20 hours ago
  • $140k - $193k

     ...single system. Combined with industry leading AI, the Motive platform gives you complete...  .... About the Role: As a Senior Software Engineer on the Equipment Monitoring team, you will...  ...firmware and IoT teams on device payloads and end-to-end data flows, making this a great... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    2 days per week

    Motive

    San Francisco, CA
    20 hours ago
  • $163k - $246.5k

     ...gets smarter as you build, with AI that learns your context to...  ....Founded in San Francisco and backed by Menlo Ventures, Felicis Ventures...  ....About the roleAs a Backend Engineer, you’ll work on our backend...  ...of the hiring process. To that end, we generate internal compensation... 
    Currently hiring
    Work at office
    Local area
    Remote work
    Weekend work
    3 days per week

    Semgrep

    San Francisco, CA
    2 days ago
  •  ...Infrastructure for the World’s Largest Dataset You'll be our seventh engineering hire. You'll have full ownership over major features, play a...  .../20 solutions. You’re great at “figuring it out”: Recall.ai is a low-structure, high-trust environment. There’s minimal... 
    Full time
    Immediate start

    Recall

    San Francisco, CA
    20 hours ago
  • $145k - $217k

    Job DescriptionSenior Full Stack Engineer, Solve VoiceJoin us at Zendesk, where we're on a mission...  ...ambition by building products rooted in AI, automation, and intelligent customer...  ...processing, and external AI/speech services.Ship end-to-end product features across backend... 
    Full time
    Remote work

    Forethought

    San Francisco, CA
    1 day ago
  • $220k - $240k

     ...OpportunityThe Forward Deployed Engineering (FDE) team tackles some of...  ...enable customers to safely adopt AI at scale.This isn't consulting...  ...deployments and feed it back to core engineering, hardening...  ...services or distributed systems with end-to-end ownership.Expertise in... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    3 days ago
  • $167k - $219k

     ...Build the Future of LearningJoin us to design and deliver AI-powered learning tools that scale across the world and unlock...  ...relevance and low latency. We're looking for a Backend Engineer who can own this end-to-end: from data ingestion and Elasticsearch index design... 
    Work at office
    2 days per week

    Quizlet

    San Francisco, CA
    4 days ago
  • $185k - $385k

    About the TeamOpenAI’s Applications Engineering organization builds and operates the products...  ...to solve a particular problemOwn problems end-to-end, and are willing to pick up whatever...  ...the job done.About OpenAIOpenAI is an AI research and deployment company dedicated... 
    Work at office
    Local area
    Worldwide
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $223k

     ...the RoleChime is looking for an Engineering Manager to lead one of the teams building Jade, our AI-powered financial assistant,...  ...trustworthy.We're looking for a Staff Software Engineer to help design...  ...reusable systems (agent loops, evals, custom tooling) that compound... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work

    Chime

    San Francisco, CA
    2 days ago
  • $140k - $190k

     ...adventure?Loft Orbital is looking for a Backend Engineer to join our Oort team. Oort is our central...  ...and technical processes, mapping them end‑to‑end—highlighting friction points and...  ...Earth observation, IoT connectivity, on-orbit AI, national security missions, and more.... 
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    2 days ago
  • $139k - $257.55k

     ...Team: You will be joining the newly formed Forward Deployment Engineering (FDE) team within Adobe’s Digital Experience organization. As part...  ...Knowledge of applied GenAI and ML with proficiency in Python AI development experience This role is eligible for bonus and equityAbout... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  • $150k - $245k

     ...the globe. Valued at US$11 billion and backed by world-leading investors including T...  ...products, and you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster...  ..., are tool-agnostic, and expect engineers to own problems end-to-end. We collaborate... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Back End Engineer, Evals - Hazel AI. Be the first to apply!