Research Scientist, Interpretability
$350kAnthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. We’re looking for researchers and engineers to join our efforts. People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks using “microscopes” we build, or as treating neural networks as binary computer programs we're trying to "reverse engineer". We aim to create a solid foundation for mechanistically understanding neural networks and making them safe (see our vision post). In the short term, we have focused on resolving the issue of "superposition" (see Toy Models of Superposition, Superposition, Memorization, and Double Descent, and our May 2023 update), which causes the computational units of the models, like neurons and attention heads, to be individually uninterpretable, and on finding ways to decompose models into more interpretable components. Our subsequent work found millions of features in Sonnet, one of our production language models, represents progress in this direction. In our most recent work, we develop methods that allow us to build circuits using features and use this circuits to understand the mechanisms associated with a model's computation and study specific examples of multi-hop reasoning, planning, and chain-of-thought faithfulness on Haiku 3.5, one of our production models. This is a stepping stone towards our overall goal of mechanistically understanding neural networks. We often collaborate with teams across Anthropic, such as Alignment Science and Societal Impacts to use our work to make Anthropic’s models safer. We also have an Interpretability Architectures project that involves collaborating with Pretraining. Responsibilities Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights Design and run robust experiments, both quickly in toy scenarios and at scale in large models Create and analyze new interpretability features and circuits to better understand how models work Build infrastructure for running experiments and visualizing results Work with colleagues to communicate results internally and publicly Qualifications Have a strong track record of scientific research (in any field), and have done some work on interpretability Enjoy team science – working collaboratively to make big discoveries Are comfortable with messy experimental science. We’re inventing the field as we work, and the first textbook is years away You view research and engineering as two sides of the same coin. Every team member writes code, designs and runs experiments, and interprets results You can clearly articulate and discuss the motivations behind your work, and teach us about what you’ve learned. You like writing up and communicating your results, even when they’re null Familiarity with Python is required for this role Location and Compensation This role is based in the San Francisco office; however, we are open to considering exceptional candidates for remote work on a case‑by‑case basis. The annual compensation range for this role is $350,000 - $850,000 USD. Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren’t able to successfully sponsor visas for every role and every candidate. If we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. Equal Employment Opportunity As set forth in Anthropic’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Safety Note We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from anthropic.com/careers email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How We're Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come Work with Us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process. #J-18808-Ljbffr Anthropic
- Radical Numerics is seeking a Member of Technical Staff, Mechanistic Interpretability, to study how multimodal genome language models represent and reason about information. This research-oriented role emphasizes model understanding, driving scientific discovery and innovation...Suggested
$216k - $270k
Scale Labs, Research Scientist — Safety Post TrainingAs the leading data and evaluation partner for frontier AI companies, Scale plays... ...Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better...SuggestedFull time- ...well known in the AI community for seminal research accomplishments at top AI labs, have run... ...a highly experienced AI Research Scientist to play a crucial role in the development... ...scalable AI pipelines to process, analyze, and interpret large-scale biological and chemical...Suggested
$225k - $300k
Research Scientist About Latent Health Healthcare today is only truly personalized for two groups: those with wealth and access, and those... ...Make and own tradeoffs between model capability, interpretability, and verifiability in high-stakes settings Collaborate with...SuggestedWork at officeImmediate start- ...pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in large-scale vision-language pretraining... ...and hallucination detection Experience with model interpretability, explainability, and uncertainty quantification in safety‑...Suggested
$140k - $200k
...powers breakthrough AI models at leading research labs and enterprises. Since 2018, we’ve... .... This is not a traditional research scientist role. You will not spend months... ...designing metrics, building eval harnesses, interpreting results critically. Familiarity with human...Work at officeFlexible hours2 days per week$200k - $400k
...species. Our goal is to tame this new fire. Goodfire is an AI interpretability research company focused on understanding and intentionally... ...Clinic, and Rakuten. The Role: We’re looking for a Research Scientist to join our team and develop new techniques for understanding...$300k - $320k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...We are seeking an exceptional Research Scientist to join our Life Sciences team at...Visa sponsorship$285k - $380k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...We are seeking a Recruiting Research Scientist to join our People Data Solutions team....Work at officeVisa sponsorshipFlexible hours- Job Title Research Scientist, Molecular Engineering Salary Not Disclosed Company Description Well-funded biotech platform startup Job Description... ...high-throughput workflows and the ability to analyze and interpret large experimental datasets. #J-18808-Ljbffr Jack & Jill
- the company is seeking a Research Scientist to advance measurable recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret results to guide R&D decisions. Senior roles exist and involve hands-on work alongside...
- Forward Deployed Research Scientist, Biology Goodfire In this role, you'll lead scientific research with partners to interpret advanced biological foundation models and develop new techniques for understanding them. Work directly with customers to conduct original research...
- Anthropic is seeking a Research Scientist to measure and understand recursive-self-improvement in large models. You will design evaluations and models, run experiments, and interpret results to guide research direction. We hire at junior and senior levels; seniors lead...Work at office
- ...pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning... ...medical knowledge representation Experience with model interpretability, explainability, and uncertainty quantification in safety‑...
- ...OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we... ...Background spanning multiple research areas (e.g., both interpretability and RL, or both systems and training methodology) Track...Immediate startFlexible hours
$350k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...the role We're looking for a Research Scientist who has done hands-on research on large...Work at officeVisa sponsorshipFlexible hours$300k - $320k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...We're seeking an exceptional Research Scientist to join the team. As a founding member...Work at officeImmediate startVisa sponsorshipFlexible hours- ...Research Scientist Position 100% on-site Position Summary The CV Translational Early Development group within the Immunology and Cardiovascular... ...Experience with wet lab experimentation, analysis & interpretation of results is required. Understanding of drug discovery...
- We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain... ....Role overview:We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion...Relocation
$120.7k - $238.6k
The OpportunityAdobe Research is looking for research scientists in Generative AI to join a world-class research team. We welcome outstanding candidates at all levels (new graduates, experienced, principal) in all related technical fields, such as Machine Learning, Deep...Full timeTemporary workLocal areaWorldwide$127k - $333.7k
THE DEPARTMENT OF MEDICINE AT THE UNIVERSITY OF CALIFORNIA SAN FRANCISCO (UCSF) is recruiting for the position of Research Scientist in the Division of Geriatrics. The appointment will be made at the level of Assistant, Associate, or Full Adjunct Professor. Candidates...- University of California, San Francisco is seeking a Professional Researcher to join the Department of Pediatrics. Appointment at Assistant... ...conducting research in the lab of an established Principal Scientist, collaborating with staff and postdocs, designing experiments,...
$218.4k - $273k
...Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings together customer-facing Researchers and... ...like Pytorch, Jax, or Tensorflow. You should also be adept at interpreting research literature and quickly turning new ideas into prototypes...Full time$216k - $270k
Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale... ...control or alignment research (e.g., scalable oversight, interpretability, debate). Experience with post‑training and RL techniques...Full time$380k
Research Scientist - Multimodal Agent, Consumer Devices | OpenAI Careers Research Scientist - Multimodal Agent, Consumer Devices Consumer... ...ensure that adaptation and personalisation remain aligned, interpretable, and bounded by clear constraints. Prototype and iterate...Work at officeImmediate startRelocation package- Scale Labs in San Francisco seeks a Research Scientist focused on Safety Post-Training to advance post-training methods and interpretability for frontier AI systems. You will design pipelines, evaluate safety properties, and help translate findings into practical guidelines...
$141.2k - $257.1k
Full Professional Researcher - Krummel labThe Krummel lab in the Department of Pathology seeks a computational scientist to help lead the UCSF Custom Immunoprofiler project (co-led... ...across patients and experimental conditions.Interpret results in biologically and clinically...$150k - $250k
...manufacturing, consumer goods, and global social organizations.We research and deploy technologies that power AI-native operations —... ...evolution. This work bridges reinforcement learning, interpretability, and meta-optimization to pioneer continuously learning enterprise...Work at office3 days per week$216.3k - $280.8k
...received.Meet the TeamAt Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance the state of... ...advanced domains such as reinforcement learning, mechanistic interpretability, and scaling laws, with applications to the cybersecurity domain...Full timeTemporary workLocal areaFlexible hours- ...out of the office. We're looking for a Senior Human Factors Researcher to join our Design Research team. This role is ideal for a researcher... ...teams reason about and apply that data in their own work, interpreting human modeling data to evaluate form factors, pressure-test...Work at officeLocal areaRemote workFlexible hours2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist, Interpretability. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA
- research associate scientist San Francisco, CA

