Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness)
$115k - $260kGEICO Insurance Agent
Why Join GEICO?
At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.
Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.
Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.
Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness)
Why Join GEICO?
GEICO is transforming how AI is built and deployed across the enterprise. As one of the largest insurers in the United States, we are investing heavily in next-generation AI platforms that empower more than 30,000 associates and enhance experiences for millions of customers.
We are looking for a Sr. Staff AI Engineer to push the frontier of what GEICO's AI Agent Platform can do, with a focus on two core systems: the agent harness that determines how agents plan, reason, use tools, and manage context, and the eval harness that tells us, with evidence, whether they are getting better. The field is moving faster than any job description can capture. Techniques that define best practice today, such as context engineering, MCP-based tool integration, skills, sub-agent orchestration, and LLM-as-judge evaluation, barely existed a few years ago and will look different a year from now. We are looking for someone who thrives in that environment and can move quickly from a new idea to a working prototype to a data-backed conclusion about whether it delivers value.
The Opportunity
The AI Agent Platform team builds use-case-agnostic infrastructure that many business workflows across GEICO build on, from claims and underwriting to internal associate tools. In this role, you'll own the intelligence layer of that platform: the agent design patterns, context strategies, and evaluation methodology that make agents on the platform reliably good at their jobs.
You'll work alongside backend platform engineers who own durable execution, scaling, and operations. Your focus is agent behavior and quality, and the science of measuring and improving it, grounded in a solid understanding of distributed systems so that what works in an experiment also works at enterprise scale. Success here comes from experimentation: forming hypotheses, designing experiments, reading traces, and iterating until performance meets the bar for real business workflows.
The model is only the starting point. Your job is to find out what actually works, prove it, and make it reusable.
What You Will Do
Agent Harness
- Design the core agent loop, including planning, reasoning, tool selection and calling, memory, and error recovery, as reusable patterns that generalize across many business workflows.
- Develop context engineering strategies: prompt architecture, context window management, summarization and compaction, memory, and retrieval (RAG) integration.
- Define how agents connect to tools and enterprise systems through the Model Context Protocol (MCP), including tool design, server patterns, and the conventions that make tools reliable for agents to use.
- Design how agents discover, load, and apply skills (packaged instructions, scripts, and resources for specific tasks) so capabilities built once can be reused across many workflows.
- Explore sub-agent and multi-agent patterns, and decide with evidence what belongs in the platform.
- Evaluate new frontier and open-weight models as they are released, and understand their tradeoffs in capability, latency, and cost for platform use.
Evaluation Harness
- Design evaluation methodology: task-level metrics, rubrics, golden datasets, LLM-as-judge graders, and human review, with automated graders calibrated against human judgment.
- Build the eval harness that lets any team define, run, and compare evaluations for their agents, spanning offline benchmarks, simulated users and environments, and online quality monitoring.
- Make eval results trustworthy by accounting for variance, statistical significance, dataset contamination, and grader bias.
- Diagnose failure modes by analyzing traces and eval results, and turn qualitative observations into measurable hypotheses.
Experiment and Iterate
- Run fast, well-designed experiments on prompts, architectures, models, and tools, owning the loop from hypothesis to result to decision.
- Iterate until agent quality meets the bar for production business workflows, and help define what that bar should be.
- Track the research frontier and industry practice, bring the best ideas into GEICO quickly, and filter out the hype.
- Share what you learn through write-ups, internal talks, reusable patterns, and playbooks that raise the level of GenAI practice across the company.
Lead
- Set technical direction for agent design and evaluation across the AI Agent Platform, in partnership with senior AI and engineering leaders.
- Mentor engineers and data scientists on GenAI engineering practice and experimental rigor.
- Partner with platform engineers, product managers, and business teams to turn workflow needs into measurable quality targets and generalized platform capabilities.
How You Work
- Intellectual curiosity. You want to understand why something works, not just whether it does.
- Speed to learn. A new model, paper, protocol, or framework drops, and you're productive with it within days.
- Adaptability. You change direction when the evidence says so, and hold strong opinions loosely.
- Empirical rigor. You trust data over intuition, and your intuition is well calibrated because of it.
- Bias to action. You'd rather build a prototype and measure it than debate it in a meeting.
Minimum Qualifications
- 10+ years of experience in software engineering, ML engineering, or applied science, with strong Python proficiency.
- Tech-lead experience: owned the design and delivery of significant systems involving multiple engineers or teams.
- Significant hands-on experience building GenAI applications and agentic systems with LLMs such as GPT, Claude, Llama, or Qwen, taken beyond prototype into real use.
- Demonstrated experience designing and running experiments to improve LLM or agent performance, including defining metrics, building evaluation sets, analyzing results, and iterating.
- Deep practical understanding of prompt and context engineering, tool calling, RAG, and agent architectures, including hands-on experience with MCP.
- Solid distributed systems fundamentals, including concurrency, fault tolerance, latency and throughput tradeoffs, and API design, with experience building or working closely with production services at scale.
- A track record of quickly learning and applying new techniques in a fast-moving field.
- Demonstrated technical leadership influencing direction across teams.
Preferred Qualifications
- Experience building evaluation frameworks or benchmarks for LLMs or agents, including LLM-as-judge methods and human annotation pipelines, using tools such as Inspect AI, Braintrust, DeepEval, or RAG.
- Experience building MCP servers and tools, or authoring skills using the Agent Skills standard (or similar packaged agent capabilities) used across teams.
- Experience with agent frameworks such as LangGraph, Microsoft Agent Framework, OpenAI Agents SDK, Claude Agent SDK, or Google ADK, and eval and observability tools such as Langfuse, LangSmith, Braintrust, or Arize Phoenix.
- Background in ML or applied science, such as statistics and experimental design, fine-tuning, reinforcement fine-tuning (e.g., RLHF, RLVR), or model evaluation.
- Experience with cloud AI platforms such as Microsoft Foundry (including Azure OpenAI and Foundry Agent Service) or Amazon Bedrock (including AgentCore).
- Publications, open-source contributions, or public writing on GenAI or agents.
- Understanding of AI security (e.g., prompt injection, tool and skill supply chain risks), safety, and responsible AI practices.
Why This Role Is Unique
- Define how agents behave, and how their quality is measured, across one of the largest insurers in the United States.
- Work hands-on with the newest models and techniques as they're released, not years later.
- See your experiments turn directly into scalable platform capabilities used by teams across the enterprise.
- Build AI capabilities that impact 30,000+ associates and help shape AI experiences used by millions of customers.
- Influence AI strategy and architecture across the enterprise, partnering with senior AI and engineering leaders on the next generation of AI at GEICO.
Annual Salary
$115,000.00 - $260,000.00The above annual salary range is a general guideline. Multiple factors are taken into consideration to arrive at the final hourly rate/ annual salary to be offered to the selected candidate. Factors include, but are not limited to, the scope and responsibilities of the role, the selected candidate’s work experience, education and training, the work location as well as market and business considerations.
GEICO will consider sponsoring a new qualified applicant for employment authorization for this position.The GEICO Pledge:
Great Company: Protecting customers through life’s twists and turns with innovation and integrity.
Great Careers: Personalized development programs, mentorship, and certification assistance.
Great Culture: Inclusive and collaborative culture rooted in shared success.
Great Rewards: Competitive pay, benefits, and flexibility to support your well-being and future.
The equal employment opportunity policy of the GEICO Companies provides for a fair and equal employment opportunity for all associates and job applicants regardless of race, color, religious creed, national origin, ancestry, age, gender, pregnancy, sexual orientation, gender identity, marital status, familial status, disability or genetic information, in compliance with applicable federal, state and local law. GEICO hires and promotes individuals solely on the basis of their qualifications for the job to be filled.
GEICO reasonably accommodates qualified individuals with disabilities to enable them to receive equal employment opportunity and/or perform the essential functions of the job, unless the accommodation would impose an undue hardship to the Company. This applies to all applicants and associates. GEICO also provides a work environment in which each associate is able to be productive and work to the best of their ability. We do not condone or tolerate an atmosphere of intimidation or harassment. We expect and require the cooperation of all associates in maintaining an atmosphere free from discrimination and harassment with mutual respect by and for all associates and applicants.
$171.9k - $300.8k
....Veza's Access Graph platform maps an organization'... ...brings together Veza's AI-native Access Graph... ...environments, and AI agents. For engineers joining Veza today, this... ...and enterprise-scale harness platform that powers... ..., re-ranking, and evaluation — as a critical dependency...PlatformSeniorWork experience placementWork at officeRemote workFlexible hours- Sr. Staff AI Engineer, Silicon Design Position Overview We are seeking a versatile Sr. Staff... ...inference optimization tools, and model evaluation harnesses. Preferred Qualifications & Skills... ...expertise in operationalizing GenAI platforms on distributed, multi‑GPU cloud...PlatformSenior
$135k - $185k
...Content Intelligence platform shaping the future... ...by advanced AI, recommendation systems... ...as an excellent engineer to join our advertising... ...systems (AI Agents) that run a closed... ...and offline/online evaluation that keep optimization... ...Hands-on with harness engineering, structured...PlatformFull timeWork experience placementLocal areaWork from home- ...Subcontracting, NO outside firms. AI Engineer – AI Agents & Generative AI Mid-to-Senior Level... ...work across multiple enterprise AI platforms. The ideal candidate combines strong... ...controls. Develop testing and evaluation processes for AI agents and RAG solutions...Platform
$180k - $250k
...Our Operational Data Platform harnesses full-census, comprehensive... ...and operations teams AI-powered insights into... ...patterns across AI agents, apps, websites, and... ...and it is not prompt engineering; the team builds... ...Familiarity with agent evaluation or LLM-ops tooling...PlatformFull time$130k - $300k
...Great Rewards, and Great Careers. Sr. Staff Software Engineer, AI Agent Platform Why Join GEICO? GEICO is... ...building several core systems: an agent harness that runs agentic workflows... ...for running offline and online evaluations, simulating scenarios, capturing...PlatformSeniorHourly payFull timeWork experience placementLocal area$227.5k - $300k
...the transformation to AI-enabled software-defined... ...motivated Senior Staff AI Engineer to join our team and... ...technical details of multi-agent orchestration, tool-... ...AI technologies, and platforms. Responsibilities:... ...engineering, and model evaluation techniques. Familiarity...PlatformSeniorWork at officeLocal areaWorldwideFlexible hoursShift work$120k - $220k
...the Content Intelligence platform shaping the future... ...information powered by advanced AI, recommendation systems... ...We're building the agent platform that powers... ...dedicated Agent Platform engineer to own this layer end-... ...eval and observability harness that runs thousands of...PlatformFull timeLocal areaWork from home$135k - $155k
...the Content Intelligence platform shaping the future... ...information powered by advanced AI, recommendation systems... ...to apply LLMs and Agent technology to real... ...work alongside senior engineers to develop AI advertising... ...safeguards, and offline/online evaluation pipelines that keep...PlatformFull timeWork experience placementInternshipLocal areaWork from home$300k
...Technical Leader Grindr is an AI-native platform powering how millions of... ...vision. This is your chance to harness cutting-edge machine... ...teams, collaborating with engineering, data science and product teams... ...ideas into tangible results. Evaluate, influence, and integrate...PlatformSeniorCasual workWork at officeImmediate startWorldwideFlexible hours$220k - $350k
...goal of enabling human life on Mars. SR. AI ENGINEER, SPECIAL PROGRAMS - TOP SECRET... ...Grok family) and government systems, platforms, and data environments. Identify pain... ...cases. Benchmark models and develop evaluation frameworks to assess performance and address...PlatformSeniorPermanent employmentFull timeTemporary workFor contractorsShift workWeekend work$175k - $287k
...business needs of the team. LinkedIn’s Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI agents and GenAI... ...for all LinkedIn AI Agents.As a Staff Engineer, you will own the end-to-end technical vision...PlatformFor contractorsWork experience placementWork at officeFlexible hours- ...Job TitleStaff Applied AI EngineerAbout Your RoleClover... ...on the Point of Sale platform. You'll define the... ....Design tool layer for agents writing to inventory —... ...on evidence, not demos.Engineer multimodal ingestion —... ...record, with real users.Evaluation-Driven Development: Specs...PlatformFull timeWork at officeMonday to FridayNight shift
- ...success. Principal AI Engineer at BairesDev In this... ...routes to specialized agents across multiple... ...retrieval strategy, and evaluation, defending your choices... ...enterprise-scale agentic AI platform. Design multi-agent... ...evolve the evaluation harness used to validate AI...PlatformLocal areaRemote workWork from homeWorldwideFlexible hours
$185k - $215k
...DescriptionIn this role, you will operate as a Staff AI Systems Engineer, setting technical direction across... ...the design, implementation, and evaluation of our Advanced Driver Assistance... ...perception, planning, control, and vehicle platform layers.Integrate automated triaging...PlatformWork experience placementLocal areaWorldwide- ...BairesDev is seeking a Principal AI Engineer to set the technical direction and own the architecture of a major enterprise AI assistant... ...You will shape the system from the ground up, ensure robust evaluation, and drive best practices in CI/CD, observability, and scalable...PlatformSeniorRemote job
- ...BairesDev is seeking an AI Engineer (Agents) to design and build autonomous, multi-step systems that can reason and act with minimal intervention... ..., implement observability, and maintain safety of agentic platforms. Strong Python skills and deep knowledge of transformer-...PlatformRemote job
$124.8k - $283.8k
...global growth with future ready, global IT platforms, applications and services. We are chartered... ...are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical evaluation and end-to-end execution of AI infrastructure...PlatformSeniorFull timeOverseasRelocation package- ...highly skilled, hands-on engineer with a passion for embedding advanced AI into the heart of... ...at integrating AI into platforms such as Microsoft AI , SAP... ...microservices, and multi-agent workflows. Your experience... ...reliability, and continuous evaluation. Mentor engineers,...PlatformSenior
$30 - $70 per hour
...Content Intelligence platform shaping the future... ...powered by advanced AI, recommendation systems... ...ranking, backstage agent workflows, multi-... ...knowledge integration, and evaluation, working with engineers and product managers... ..., evaluation harnesses). Understanding of...PlatformFull timeFor contractorsInternshipLocal areaWork from home$205.5k - $278k
...financial technology platform that powers... ...a Senior Staff Product Manager... ...and roadmap for Agent Runtime capabilities... ...foundational AI runtime that enables... ...rapidly build, evaluate, and operate... ...with AI, Engineering, and Experience... ...including agent harnesses, durable execution...PlatformSeniorWork at officeWorldwide3 days per week$126k - $204.5k
...Integrity, and Inclusion. We weave AI into the fabric of everything we... .... Partner closely with engineering, product, and infrastructure teams to evaluate feature requests and translate business... ...operating microservices on containerized platforms (Docker, Kubernetes, AWS EKS, GKE...PlatformSeniorFull timeWork at officeVisa sponsorshipWork visa- ...and cost. Implement evaluation and continuous... ...offline and online eval harnesses, golden sets, human-in... ...tools: alarms and KPI platforms, ticketing, inventory/... ...functionally with data engineering, platform/security, and... ...vector stores (e.g., Azure AI Search) Prompt...PlatformPermanent employmentContract workLocal area
$170k - $185k
...all continents. Our leading AI technology is the backbone of... ...the first enterprise-grade AI Agent platform built for mission-critical applications... ...As a Senior AI/ML Engineer, you will design, develop, and... ...orchestration frameworks. Evaluate and integrate emerging AI technologies...PlatformSeniorFor contractorsFlexible hours$108k - $170k
About Us Observe.AI is the AI Agents platform for customer experience, designed to help organizations... ...Us We're looking for an AI Agent Engineer to lead the charge in building and deploying... ...integrations, telephony setup, and evaluation frameworks. Client Engagement: Act...PlatformFull timeWork at officeLocal areaRemote workFlexible hours- ...Description Senior Software Engineer Job Type: Contractor (~15 hours... ...Engineers to support an AI training project by creating reinforcement... ...learning environments that evaluate AI models on complex software... ...solutions. Evaluate AI agents' ability to reason through...Remote jobFor contractors
$262k - $364k
...developers to test complex agent behaviors, tool... ...reproducible evaluation infrastructure and automated test harnesses capable of... ...Design and scale the platform infrastructure that... ...Build the metrics engines and dashboards to... ...most comprehensive AI platform in the industry...PlatformSenior- ...HIRING: Senior / Staff Full Stack Engineer – AI Location: Mountain View, CA – 3... ...team building an AI-powered platform that creates real, data-... ...LangGraph / LangChain / Claude Agent SDK ✅ AWS – EKS, EC2, S3... ...⭐ Good to Have: * AI Evaluation platforms/tools * Active...PlatformSeniorTemporary work
$175k - $265k
...the potential of generative AI to power the transformation of... ...infrastructure layer that every engineering team and customer depends on... ...cloud environments, and the platform services used to deploy and... ...maintain a consistent and fair evaluation of al applicants. Thank you...PlatformSenior- ...Role We are seeking a Senior Data / AI / ML Software Engineer with 7+ years of experience building... ...enjoys designing and improving core platform components at the intersection of software... ..., analyze, and derive insights Evaluate system quality and performance to continuously...PlatformSeniorFull timeContract workInternship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness). Be the first to apply!
- technology administrator Palo Alto, CA
- assistant engineer Palo Alto, CA
- staff engineer Palo Alto, CA
- senior staff systems engineer Palo Alto, CA
- engineering aide Palo Alto, CA
- senior ai engineer Palo Alto, CA
- ai developer Palo Alto, CA
- ai prompt engineer Palo Alto, CA
- ai engineer Palo Alto, CA
- ai engineer remote Palo Alto, CA




