FAR.AI
Berkeley, CA
Research Scientist, Applied White-Box Methods
We tailor your resume to this role and apply for you in seconds.
Apply to Research Scientist, Applied White-Box Methods at FAR.AIJob details
- Location
- Berkeley, CA
- Work type
- Onsite
- Compensation
- $150,000 - $250,000/yr
- Visa
- Sponsorship available
- Posted
- 3 days ago
- Apply on
- jobs.ashbyhq.com
About this role
FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. The Research Scientist will advance the Applied White-Box Methods team’s research agenda by developing, evaluating, and applying methods that use model internals to improve AI safety, publishing findings, and engaging with the AI alignment community.
What you'll do:
- As a Research Scientist on the team, you will take ownership of and accelerate the team's research agenda, publish findings broadly, and engage with the AI alignment community
- You are encouraged to propose new directions within the team's agenda
- You are welcome and encouraged to attend relevant conferences and other community events, and leverage FAR.AI’s existing comprehensive infrastructure for events convening and government relations as you see fit
- Beyond FAR.AI, you can work with national AI safety institutes, frontier model developers, and top academics
What they're looking for:
- Two backgrounds are particularly well suited to this role: researchers with some experience in more fundamental interpretability who want to evaluate and refine those methods in realistic settings, including agentic coding, long contexts, and realistic threat models; and researchers from evaluations, AI control, red-teaming, or reinforcement learning who have begun working with model internals and want to deepen that work
- If you are new to LLM research, that is also ok! But please be prepared to explain how your previous research experience (e.g. in applied ML) could be leveraged for our work, and how you are engaging with the field of technical AI safety research today
- Hands-on experience applying at least one white-box method to a real model (activation explainers, SAEs, steering, attribution, influence functions, probes, or similar), and an informed view of its limitations
- A track record in AI safety: a paper, a fellowship project, or substantive public writing
- Experience with evaluations, AI control, red-teaming, reinforcement learning, or post-training of LLMs
- The ability to communicate novel methods and results clearly to technical and non-technical audiences
- A PhD or several years of research experience in computer science, machine learning, physics, statistics, or a related field
- Previous experience in applied ML for other fields: e.g. biology, chemistry, materials science, robotics, etc
- Location: Berkeley, CA. This role is in-person; we strongly prefer candidates who are in the Bay Area or willing to relocate, and we sponsor visas for in-person employees
- Hours: full-time (40 hours/week)
- In both cases, prior hands-on experience with white-box methods and a demonstrated interest in their practical application are preferred
- Alumni of programs such as MATS, Astra, Anthropic Fellows, SPAR, Pivotal, LASR or similar programs are especially encouraged to apply
Benefits:
- We sponsor visas for in-person employees.
- Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date
- Retirement - 401(k) plan with up to 2% match
- PTO - 25 days Paid Time Off per year, accrued weekly and up to 10 days of paid sick leave per year
- Paid Leave - Paid Bereavement, Family, Medical and Pregnancy Disability Leave
- WFH Stipend & Equipment - Work computer and stipend provided for eligible employees
- Catered Meals (Berkeley Office Only) - Catered lunches and dinners on workdays at our office
- Paid work trial of up to one week
Ready to apply to FAR.AI?
ApplyBolt finds matching jobs, tailors your resume, and submits applications for you.