Research Scientist, Interpretability
$350kAnthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. We’re looking for researchers and engineers to join our efforts. People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks using “microscopes” we build, or as treating neural networks as binary computer programs we're trying to "reverse engineer". We aim to create a solid foundation for mechanistically understanding neural networks and making them safe (see our vision post). In the short term, we have focused on resolving the issue of "superposition" (see Toy Models of Superposition, Superposition, Memorization, and Double Descent, and our May 2023 update), which causes the computational units of the models, like neurons and attention heads, to be individually uninterpretable, and on finding ways to decompose models into more interpretable components. Our subsequent work found millions of features in Sonnet, one of our production language models, represents progress in this direction. In our most recent work, we develop methods that allow us to build circuits using features and use this circuits to understand the mechanisms associated with a model's computation and study specific examples of multi-hop reasoning, planning, and chain-of-thought faithfulness on Haiku 3.5, one of our production models. This is a stepping stone towards our overall goal of mechanistically understanding neural networks. We often collaborate with teams across Anthropic, such as Alignment Science and Societal Impacts to use our work to make Anthropic’s models safer. We also have an Interpretability Architectures project that involves collaborating with Pretraining. Responsibilities Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights Design and run robust experiments, both quickly in toy scenarios and at scale in large models Create and analyze new interpretability features and circuits to better understand how models work Build infrastructure for running experiments and visualizing results Work with colleagues to communicate results internally and publicly Qualifications Have a strong track record of scientific research (in any field), and have done some work on interpretability Enjoy team science – working collaboratively to make big discoveries Are comfortable with messy experimental science. We’re inventing the field as we work, and the first textbook is years away You view research and engineering as two sides of the same coin. Every team member writes code, designs and runs experiments, and interprets results You can clearly articulate and discuss the motivations behind your work, and teach us about what you’ve learned. You like writing up and communicating your results, even when they’re null Familiarity with Python is required for this role Location and Compensation This role is based in the San Francisco office; however, we are open to considering exceptional candidates for remote work on a case‑by‑case basis. The annual compensation range for this role is $350,000 - $850,000 USD. Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren’t able to successfully sponsor visas for every role and every candidate. If we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. Equal Employment Opportunity As set forth in Anthropic’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Safety Note We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from anthropic.com/careers email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How We're Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come Work with Us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process. #J-18808-Ljbffr Anthropic
$225k - $300k
...Research Scientist About Latent Health Healthcare today is only truly personalized for two groups: those with wealth and access... ...Make and own tradeoffs between model capability, interpretability, and verifiability in high-stakes settings Collaborate...SuggestedWork at officeImmediate start$200k - $400k
...Research Scientist San Francisco, CA About Goodfire Goodfire is a research company using interpretability to understand, learn from, and design AI systems. Our mission is to build the next generation of safe and powerful AI—not by scaling alone, but by understanding...SuggestedRemote work$216k - $270k
Scale Labs, Research Scientist — Safety Post TrainingAs the leading data and evaluation partner for frontier AI companies, Scale plays... ...Training you will develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better...SuggestedFull time- ...pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in large-scale vision-language pretraining... ...and hallucination detection Experience with model interpretability, explainability, and uncertainty quantification in safety‑...Suggested
$140k - $200k
...powers breakthrough AI models at leading research labs and enterprises. Since 2018, we’ve... .... This is not a traditional research scientist role. You will not spend months... ...designing metrics, building eval harnesses, interpreting results critically. Familiarity with human...SuggestedWork at officeFlexible hours2 days per week- ...well known in the AI community for seminal research accomplishments at top AI labs, have run... ...a highly experienced AI Research Scientist to play a crucial role in the development... ...scalable AI pipelines to process, analyze, and interpret large-scale biological and chemical...
$147k - $210k
Drive projects by defining key research questions.Design, implement, and evaluate experiments... ...with training, evaluating, and interpreting large language models.Experience working... ...simulation.We are looking for a Research Scientist to develop cutting-edge social...$300k - $320k
...Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...We're seeking an exceptional Research Scientist to join our Life Sciences team at Anthropic...Work at officeVisa sponsorshipFlexible hours$230k - $400k
...Description Job Description Our mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and... ...as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together...Work experience placementWork at officeHome officeVisa sponsorshipRelocation packageFlexible hours- ...pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in vision foundation models to join our... ...language models and multimodal learning Experience with model interpretability and explainability methods Understanding of clinical...
- Forward Deployed Research Scientist, Biology Goodfire In this role, you'll lead scientific research with partners to interpret advanced biological foundation models and develop new techniques for understanding them. Work directly with customers to conduct original research...
- ...OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we... ...Background spanning multiple research areas (e.g., both interpretability and RL, or both systems and training methodology) Track...Immediate startFlexible hours
- ...pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning... ...medical knowledge representation Experience with model interpretability, explainability, and uncertainty quantification in safety‑...
$245k - $285k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and... ...the Role We are looking for biological scientists to help build safety and oversight mechanisms...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- Anthropic in San Francisco seeks a Research Scientist to measure recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret results to guide research direction. Senior candidates will combine hands-on work with...
$350k
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...the role We're looking for a Research Scientist who has done hands-on research on large...Work at officeVisa sponsorshipFlexible hours$285k - $380k
...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI... ...a quickly growing group of committed researchers, engineers, policy experts, and business... ...We are seeking a Recruiting Research Scientist to join our People Data Solutions team....Work at officeVisa sponsorshipFlexible hours$220k - $315k
Forward Deployed Research Scientist (Biology) About Goodfire Behind our name: Like fire, AI holds the potential for both immense benefit... .... Our goal is to tame this new fire. Goodfire is an AI interpretability research company focused on understanding and intentionally...- Anthropic is seeking a Research Scientist to measure and understand recursive-self-improvement in large models. You will design evaluations and models, run experiments, and interpret results to guide research direction. We hire at junior and senior levels; seniors lead...Work at office
- ...Immunology and Cardiovascular Thematic Research Center (ICV-TRC) at Client seeks to understand... ...talented and highly motivated research scientist interested in a fast-paced scientific... ...wet lab experimentation, analysis & interpretation of results is required. •...
- University of California, San Francisco is seeking a Professional Researcher to join the Department of Pediatrics. Appointment at Assistant... ...conducting research in the lab of an established Principal Scientist, collaborating with staff and postdocs, designing experiments,...
$120.7k - $238.6k
The OpportunityAdobe Research is looking for research scientists in Generative AI to join a world-class research team. We welcome outstanding candidates at all levels (new graduates, experienced, principal) in all related technical fields, such as Machine Learning, Deep...Full timeTemporary workLocal areaWorldwide$127k - $333.7k
THE DEPARTMENT OF MEDICINE AT THE UNIVERSITY OF CALIFORNIA SAN FRANCISCO (UCSF) is recruiting for the position of Research Scientist in the Division of Geriatrics. The appointment will be made at the level of Assistant, Associate, or Full Adjunct Professor. Candidates...- We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain... ....Role overview:We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion...Relocation
$120k - $150k
Senior Bioanalytical Scientist - Permanent - San Francisco, CALead the science behind tomorrow's medicines by advancing bioanalytical... ...and non-clinical studies such as sample analysis and data interpretation using appropriate bioanalytical technologies and instrumentationOperate...Permanent employment$380k
Research Scientist - Multimodal Agent, Consumer Devices | OpenAI Careers Research Scientist - Multimodal Agent, Consumer Devices Consumer... ...ensure that adaptation and personalisation remain aligned, interpretable, and bounded by clear constraints. Prototype and iterate...Work at officeImmediate startRelocation package$300k - $320k
Research Scientist, Life Sciences (Chemistry) About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed...Contract workWork at officeVisa sponsorshipFlexible hours$216k - $270k
Research Scientist, AI Controls and Monitoring Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation... ...control or alignment research (e.g., scalable oversight, interpretability, debate). Experience with post-training and RL techniques...Full time$350k
...Research Engineer / Scientist, AlignmentSan Francisco, CAAbout AnthropicAnthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group...Work at officeVisa sponsorshipFlexible hours- ...and accelerate delivery across the company. We also build the research platform that powers AI at scale and explore next-generation GenAI... ..., and continuous monitoring benchmarks. In addition to interpreting existing standards, you will advocate for best practices, anticipate...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist, Interpretability. Be the first to apply!
- materials scientist San Francisco, CA
- scientist assay development San Francisco, CA
- entry level research scientist San Francisco, CA
- health scientist San Francisco, CA
- quality control scientist San Francisco, CA
- deep learning scientist San Francisco, CA
- research associate scientist San Francisco, CA
- application scientist San Francisco, CA
- scientist antibody discovery San Francisco, CA
- senior analytical scientist San Francisco, CA


