
Anthropic Launches $5M Grant Program for AI Wellbeing Evaluations

Anthropic announced a $5 million grant program to fund independent research on how AI affects user wellbeing. Grantees receive direct funding, model access, and technical support, and are expected to publish open-source evaluations that any developer can use; they work independently. The program responds to the fact that AI systems are now conversational partners and sources of emotional support, while the industry lacks clear standards for how models should behave when users seek companionship or navigate a mental health crisis. Anthropic argues wellbeing is hard to evaluate because a single response may look accurate but require more context: risk can emerge over a long conversation, and the same response can be harmful in a different context, such as diet advice given to someone with a history of disordered eating.
The announcement includes guidance from Anthropic‘s Safeguards team on what makes a wellbeing evaluation rigorous. Evaluations should state what they measure and why it matters; involve clinical and subject-matter experts in design and validation; test both precautions and harms, meaning the risks of overcompliance and overrefusal; reflect actual usage patterns, often through multi-turn scenarios where risk escalates and context shifts; and validate graders against real subject-matter experts. Anthropic says the right approach will need to evolve with models and usage, and the program invites clinicians, psychologists, methodologists, and others to contribute. Applications are due September 21, and selected applicants are notified October 5.


