Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Researcher, Recursive Self-Improvement Safety

OpenAI

About the team Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. Mitigating the frontier risks resulting from these capabilities is paramount to OpenAI’s ability to continue deploying models safely. The Preparedness team is dedicated to addressing these critical risks. Our work includes: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misalignment safeguards, alignment tools, and on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework, and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role Preparedness is hiring strong technical executors to support preparations for accelerated AI development, which may culminate in recursive self-improvement. This work relies on anticipating misalignment risks that might exist in the future, but might not exist now; so it’s especially important that people in this role are tasteful and strategic. The role is wide-ranging, covering any mitigation for loss of control risk, spanning the design and implementation of better pre-deployment risk-assessment, control measures, RSI-relevant training interventions, and turning one’s technical work into established institutional practices and external-facing communications. Below is a subset of our focus areas: Scalable oversight: Establishing practices for model misbehavior monitoring and oversight which remain effective in superhuman model capability regimes, with a focus on bridging from today’s monitoring approaches to future-proof ones. Automated auditing: As model capabilities increase, we’ll increasingly rely on automated approaches for finding the most severe forms of model misalignments. We’ll both need to sift through large swaths of production traffic to find the most egregious misalignments, and reliably elicit tail risks before deployment. Rigorous monitorability: Rigorous testing and red-teaming of our measurements of model misbehavior related to loss-of-control (e.g. reward hacking, sandbagging, scheming). This includes better understanding monitorability, and e.g. preparing for potential losses of Chain-of-Thought monitorability. Model behavior science : Design experiments and evaluations to understand the extent to which models are problematically misaligned, or their safety-relevant capabilities lag behind dangerous capabilities. This may include training model organisms of misbehavior for behaviors not currently present in production, or training interventions to increase safety-relevant capabilities. Coordination and verification : Prototype technical mechanisms for verifying compliance with potential future AI safety agreements. AI R&D risk measurement: Track progress toward automation of technical staff to inform OpenAI’s near-term investments in alignment and security. Maintaining and strengthening RSI safety cases: We’re especially interested in identifying and addressing blindspots of mitigation areas which we may have missed. Generally, our team alternates between performing rigorous hypothesis-driven research and turning our insights into interventions or control systems which impact production models, with occasional support of engineering teams. In this role, you will: Carefully consider the problems OpenAI might face in the future and how to prepare for them. Turn an open-ended objective like “prepare for future misalignment threats” into a much more concrete direction (e.g. “stress-test monitors for scheming”) – prioritizing the work that is most useful to start right now. Execute quickly, building scrappy prototypes, and then improving them iteratively until they become established components of our safety pipelines. Secure buy-in from other staff at OpenAI when necessary, and communicate your work clearly. Collaborate with or manage other staff as needed, since we might need to rapidly scale to tackle these problems quickly. You might thrive in this role if you: Are an exceptional technical executor. Have strong strategic and research taste: you can prioritize effectively in domains with weak feedback loops. Are passionate about mitigating the risks associated with recursive self-improvement. Are driven by a desire to do whatever work most positively impacts the future of AI development. Bonus: you have already done work in one of the domains listed above (ML research, AI alignment, AI verification etc). About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Aff… Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology. #J-18808-Ljbffr OpenAI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Researcher, Recursive Self-Improvement Safety in San Francisco, CA vacancy
  •  ...initiatives in accelerating AI development while mitigating risks from recursive self-improvement. The role emphasizes strategic thinking, rapid prototyping, and turning insights into scalable safety pipelines. You'll collaborate across teams to stress-test monitors,... 
    Suggested

    OpenAI

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...goods, and global social organizations.We research and deploy technologies that power AI-...  ...Distyl itself. Our work spans research into self-constructing systems, the development of...  ...don’t just want to drive incremental improvements on benchmarks or optimize an existing... 
    Suggested
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    2 days ago
  • Distyl AI is seeking researchers to push the envelope of enterprise AI, building systems that redefine how software is used rather than incremental improvements. You’ll shape self-improving architectures that learn from their performance and adapt in real time. Our team... 
    Suggested
    Work at office

    Distyl

    San Francisco, CA
    4 days ago
  • Anthropic in San Francisco seeks a Research Scientist to measure recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret results to guide research direction. Senior candidates will combine hands-on work... 
    Suggested

    Anthropic

    San Francisco, CA
    4 days ago
  • Anthropic is seeking a Research Scientist to measure and understand recursive-self-improvement in large models. You will design evaluations and models, run experiments, and interpret results to guide research direction. We hire at junior and senior levels; seniors lead... 
    Suggested
    Work at office

    Anthropic Limited

    San Francisco, CA
    2 days ago
  • $295k

    About the TeamThe Safety Systems team is responsible for various safety work to ensure...  ...trust and transparency.The Model Safety Research team aims to fundamentally advance our capabilities...  ...s core model training and launch safety improvements in OpenAI’s products.Set the research... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    6 days ago
  • Distyl in San Francisco is seeking an Applied AI Researcher for the System Self-Construction team. You will design architectures enabling autonomous generation and refinement of sub-systems, pushing the frontier of self-constructing AI. Hybrid in-office collaboration is... 
    Work at office
    3 days per week

    SupportFinity™

    San Francisco, CA
    1 day ago
  • ABOUT THE COMPANY We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site. ABOUT THE ROLE You'll lead our work on model... 

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  •  ...appropriately. The Chat and Multimodal Safety team is responsible for ensuring that OpenAI...  ...these experiences. We develop the research, training methods, and evaluations needed...  ...scaling, and inference tradeoffs. Have improved frontier model behavior through post‑training... 
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    4 days ago
  • ABOUT THE COMPANY We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site ABOUT THE ROLE You'll be researching making... 
    Shift work

    MLSys 2020

    San Francisco, CA
    3 days ago
  • About the Company We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site. About the Role You'll be researching the agents... 
    Shift work

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  •  ...advance. We believe this can be dramatically improved. At Axiom, we generate and curate...  .... We are looking for a machine learning researcher to help define and build the core AI systems...  ...systems. Work on contrastive learning, self-supervised learning, semi-supervised... 

    axiombio

    San Francisco, CA
    1 day ago
  • $218.7k - $249.6k

    Applied Researcher I Overview: At Capital One, we are creating trustworthy and reliable AI...  ...and work with stakeholders to identify and improve the status quo. You’re passionate about...  ...of the following: training optimization, self-supervised learning, robustness, explainability... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we...  ...and work with stakeholders to identify and improve the status quo. You’re passionate about...  ...of the following: training optimization, self-supervised learning, robustness, explainability... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy...  ...and work with stakeholders to identify and improve the status quo. You’re passionate about...  ...of the following: training optimization, self-supervised learning, robustness, explainability... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  •  ...growing team comprised of quantitative researchers, software engineers, product managers, designers...  ...costs, continuously researching improvements to our optimization formulations and ensuring...  ...in an object oriented language. A self-starter who embraces ownership and... 
    Work at office
    Visa sponsorship
    Flexible hours

    Frec

    San Francisco, CA
    1 day ago
  •  ...The San Jose State University Research Foundation (SJSURF) is...  ...everyone can show up as their own self and have an opportunity to contribute...  ...that airspace operations and safety activities continue to...  ...propose HAT Lab operational improvements. The HAT Lab, like NASA, prizes... 
    Remote work
    Relocation
    3 days per week

    San Jose State University Research Foundation

    San Francisco, CA
    2 days ago
  • OpenAI in San Francisco seeks exceptional researchers to push the frontier of safety mitigations, helping derisk frontier models and advance techniques from interpretability, robustness and alignment to ensure safe deployments. This role requires deep technical experience... 

    Triwill Group

    San Francisco, CA
    3 days ago
  • $295k

    About the team The Safety Systems org ensures that OpenAI's most capable models can be responsibly developed and deployed. We build evaluations...  ...productivity can also accelerate exploitation. As a Researcher for cybersecurity risks, you will help design and implement an... 

    Slope

    San Francisco, CA
    10 hours ago
  •  ...customers almost immediately. No speculative research track here. If you want your work to hit...  ...benchmarks What You'll Bring A genuine, self-driven interest in this area of research,...  ...to do Hands-on experience building or improving TTS, ASR, speech-to-speech, or neural audio... 
    Permanent employment
    Full time
    Immediate start

    DeepRec.ai

    San Francisco, CA
    2 days ago
  • $293k - $405k

    About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities... 

    Neura Market

    San Francisco, CA
    4 days ago
  • $100k - $150k

     ...interaction. We are a team of passionate researchers and engineers developing technology that...  ..., and non-stationary signal modeling to improve robustness across users, sessions, and...  ...space models, diffusion/denoising models, self-supervised learning) for time-series and... 
    Work at office

    AXION

    San Francisco, CA
    3 days ago
  • $250k

     ...Transluce is a fast-moving nonprofit research lab building the public tech stack for...  ...companions impact mental health and child safety, and we’re improving outcomes for millions of sensitive AI...  ..., good experimental design, epistemic self-awareness and transparency. Ability... 
    Visa sponsorship

    Transluce

    San Francisco, CA
    1 day ago
  • Sharpe Search is seeking an elite quantitative researcher to build a self-improving hedge fund powered by AI-driven models. The candidate will own the end-to-end research lifecycle, from idea generation to production, with significant equity and extreme ownership. Ideal... 

    Sharpe Search

    San Francisco, CA
    4 days ago
  • OpenAI is looking for a Senior Researcher focused on AI safety, located in San Francisco. This role involves setting research directions for ensuring safe AGI, developing AI monitor models, and collaborating with cross-functional teams to uphold safety standards. The ideal... 

    OpenAI

    San Francisco, CA
    3 days ago
  • OpenAI in San Francisco is seeking exceptional researchers to push the frontier of safety mitigations for deployed models. You will help derisk frontier models by developing novel safety mitigations and applying methods from interpretability, control, and alignment to... 

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...of the grid that affect reliability and safety. Gridware’s advanced Active Grid Response...  ...mitigation. This comprehensive approach helps improve safety, reduce outages, and ensure the...  ...is Gridware's first dedicated design research role, and it is a foundational hire for our... 
    Local area
    Shift work

    Gridware

    San Francisco, CA
    20 days ago
  • $405k

    Staff+ Researcher, Cybersecurity Products About Anthropic Anthropic’s...  ...generations, and they continue to improve with each release. That...  ...applications Familiarity with the safety considerations of AI in...  ...selecting “Yes.” Voluntary Self-Identification For government... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic Limited

    San Francisco, CA
    1 day ago
  • $380k

    About the TeamThe Agent Safety team works to ensure that increasingly capable AI agents...  ...re looking for a safety&security minded researcher or engineer who can reason rigorously about...  ...tool use, and other harmful outcomes.Improve the safety-productivity tradeoff by measuring... 
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  • OpenAI is seeking an experienced security researcher to help mitigate AI threats and safeguard systems as AI agents become more capable. The role focuses on designing robust defenses and coordinating with teams to maintain safeguards across our platform. You will identify... 

    Neura Market

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Researcher, Recursive Self-Improvement Safety. Be the first to apply!