
Vals raises $40M for private, real-world AI benchmarking

Vals, an AI benchmarking startup founded in 2024, wants to replace legacy evaluation systems that AI companies have learned to game. The company recently raised a $40 million Series A led by Andreessen Horowitz, following a seed round led by 8VC and Bloomberg Beta. Co-founder Rayan Krishnan, a 25-year-old who previously interned at Palantir and worked at Microsoft and Stanford’s AI lab, says he started Vals after seeing capable new models arrive faster than academic benchmarks could measure them. With AI being integrated across society, he argues, benchmarks should verify that models actually do what companies advertise.
Vals differentiates itself by keeping its test materials private, unlike many public benchmarks that companies can train against and effectively cheat on. Instead of measuring general intelligence with exam-style questions, Vals evaluates models on complex, industry-specific tasks in law, finance, and coding. Krishnan says the goal is to look at real impacts: whether models can produce work at human quality within a domain, and also what negative consequences could arise if models ‘ran wild in the world.’ The startup is expanding into areas including recursive self-improvement, mental health, cybersecurity, biosecurity, and applying the Geneva Convention in armed conflict.
Companies pay Vals to test their models, a model Krishnan compares to students paying the College Board for the SAT. The evaluations help vendors troubleshoot and improve, and are increasingly used by buyers deciding which AI models to acquire. Vals reports that revenue is eight times what it was last year. The team has grown from eight at the start of the year to 25, with plans to add 10 to 15 more employees and move to a larger office. It also recently launched a program to provide model evaluations to federal agencies.
Krishnan sees this kind of benchmarking as central to the AI industry’s next phase. As AI companies approach public markets, he expects evaluations to drive model usage, appear in public filings, and shape investment decisions, naming Anthropic and OpenAI as companies that could be public soon. Vals‘ private, domain-focused approach is an attempt to make benchmarks not just a PR tool but a reliable check on what models can actually do.


