Vals AI’s $40M Raise Signals a New Standard: Who Audits the AI Auditors?

CryptoFox Guide

Speed reveals truth; patience reveals value.

The narrative that AI models are black boxes is no longer a fringe concern. Last week, Vals AI closed a $40 million Series A round led by a16z at a $400 million valuation, positioning itself as the first independent, third-party evaluation layer for large language models (LLMs). The pitch is seductive: instead of relying on static benchmarks like GSM8K or HumanEval—which are increasingly contaminated by training data—Vals claims to assess models on real-world tasks extracted from GitHub pull requests. The hook? OpenAI, Anthropic, Google, Meta, and xAI have all allegedly cited Vals’s results in their model cards.

Vals AI’s $40M Raise Signals a New Standard: Who Audits the AI Auditors?

Context: Why Now? The AI evaluation market has been a silent war. For years, model performance has been measured by academic benchmarks that are gamed, leaked, or simply irrelevant to production environments. The industry’s dirty secret: every major model vendor optimizes for these benchmarks, creating a “Goodhart’s Law” feedback loop where the metric becomes the target. Enterprise buyers, meanwhile, struggle to compare models on their own codebases or workflows. Vals enters this void with a seemingly simple solution: extract hidden tests from historical pull requests across any GitHub repository, run them against the model, and grade the output automatically. The “private, personalized” evaluation is the core product—and the entry point for B2B sales.

But the timing is non-trivial. The post-Dencun era in crypto, for example, has shown that infrastructure layers can become saturated quickly. Similarly, the AI evaluation layer is still nascent, but capital is flooding in. a16z’s bet on Vals is not a bet on current revenue (which remains undisclosed) but on the belief that third-party evaluation will become a mandatory component of AI procurement, much like SOC 2 audits are for cloud services. The valuation of $400 million for a company with no public ARR and a vague “8x revenue growth” claim (likely a misstatement of “8x vs. 2025 annual projection”) signals that VCs are buying the category, not the company.

Core: Key Facts and Immediate Impact

Let’s dissect the tech. Vals’s innovation is not in model architecture but in evaluation infrastructure. They pull actual developer tasks from any public GitHub repository’s history—issues, pull requests, code reviews—and convert them into hidden tests. The model is asked to complete the task, and the output is compared against the ground truth (the actual merged code). This is a direct productization of the SWE-bench paradigm, but with a twist: it’s cross-domain, covering finance, law, and medicine. The immediate impact is twofold:

  1. Enterprise adoption: Companies can now evaluate models against their own private codebases, lowering the switching cost between providers. This shifts decision-making from “which model is best on the leaderboard” to “which model is best for my specific repo.”
  2. Model accountability: If major labs cite Vals’s results, it creates a de facto standard. This could force smaller open-source models to also undergo third-party evaluation, accelerating the “audit economy.”

Based on my experience reverse-engineering 0x V2’s contract architecture in 2017, I know that speed of verification is often more important than depth of analysis. Vals’s approach is fast, but it inherits the same risks as any black-box testing: the test set itself can be contaminated, and the model’s performance on “hidden” tasks may not generalize to unseen workflows.

From a data perspective, Vals claims to have evaluated models across thousands of repositories. But the sample size and statistical significance of these evaluations are not disclosed. In my analysis of Aavegotchi’s NFT data in 2021, I learned that small sample sizes can produce misleading narratives. The same applies here: if Vals’s evaluations are based on a narrow set of repositories (e.g., only popular open-source projects), the results may not be representative of production environments.

Contrarian: The Unreported Angle

Here’s the problem no one is talking about: Vals’s independence is compromised by its business model. The same companies that are being evaluated—OpenAI, Anthropic, etc.—are also potential customers. If Vals charges model vendors for “evaluation as a service,” the incentive to produce favorable reports is real. The article does not disclose whether the model card citations are paid or unpaid. If they are paid, the whole “third-party” claim is a farce. It’s the same conflict of interest we see in the crypto space with audit firms that are paid by the same protocols they audit. The Terra/Luna collapse taught us that paid audits can be selective.

Vals AI’s $40M Raise Signals a New Standard: Who Audits the AI Auditors?

Moreover, the technical risk of reverse-engineering the hidden tests is substantial. If Vals uses public GitHub history, model vendors can train on that data. Vals claims to use private repositories for enterprise customers, but that creates a two-tier system: public benchmarks vs. private ones. The private ones are more trustworthy, but they also require deep integration with the customer’s codebase, which is a high-friction sell.

Another blind spot: cross-domain evaluation. Finance, law, and medicine require domain-specific verification. Can Vals’s automated tests accurately assess a legal contract’s compliance or a medical diagnosis’s correctness? Unlikely without human-in-the-loop validation. The article does not mention the human cost of labeling these tasks, which suggests the company may be over-promising on automation.

Finally, the article’s source is a blockchain monitoring channel called “Dongcha Beating.” This is a red flag: the information is drawn from an anonymous, non-credentialed source. In my 2024 Bitcoin ETF whitepaper breakdown, I learned that institutional-grade analysis requires verifiable sources. The lack of independent verification for Vals’s revenue claims, customer count, and model card citations means the entire narrative rests on trust. In a market that is already skeptical of AI hype, that trust is fragile.

Takeaway: What to Watch Next

The real test will be whether Vals can resist the gravitational pull of its own investors. a16z has a portfolio of AI companies; if Vals starts auditing those companies, the conflict of interest will be hard to ignore. Watch for three signals: - Disclosure of pricing and customer contracts: If Vals publishes a transparent fee structure for model vendors, it will increase credibility. - Independent third-party verification of its own evaluation methodology: Who audits the auditors? If Vals submits its own tests to a public audit, it would set a standard. - Adoption in regulated industries: If financial regulators or healthcare bodies cite Vals’s evaluations as evidence of model safety, the market will legitimize. Until then, this is a high-stakes gamble on the “audit economy” that may or may not materialize.

Speed reveals truth; patience reveals value. The truth of Vals’s claims will emerge not in press releases, but in the next model card update that either includes or excludes their logo. I’ll be watching the git history.

Market Prices

BTC Bitcoin
$62,992.6 +0.33%
ETH Ethereum
$1,879.32 +0.30%
SOL Solana
$75.19 -0.63%
BNB BNB Chain
$611.6 +0.58%
XRP XRP Ledger
$1 -0.02%
DOGE Dogecoin
$0.0701 +0.59%
ADA Cardano
$0.1792 -1.70%
AVAX Avalanche
$6.59 +3.53%
DOT Polkadot
$0.7777 +3.01%
LINK Chainlink
$9.26 +5.42%

Fear & Greed

34

Fear

Market Sentiment

Event Calendar

{{年份}}
15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

28
03
unlock Arbitrum Token Unlock

92 million ARB released

12
05
halving BCH Halving

Block reward halving event

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

18
03
unlock Sui Token Unlock

Team and early investor shares released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

Market Cap

All →
1
Bitcoin
BTC
$62,992.6
1
Ethereum
ETH
$1,879.32
1
Solana
SOL
$75.19
1
BNB Chain
BNB
$611.6
1
XRP Ledger
XRP
$1
1
Dogecoin
DOGE
$0.0701
1
Cardano
ADA
$0.1792
1
Avalanche
AVAX
$6.59
1
Polkadot
DOT
$0.7777
1
Chainlink
LINK
$9.26

Tools

All →

Altseason Index

44

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

🐋 Whale Tracker

🔵
0x2a1d...39de
1h ago
Stake
6,075 SOL
🔴
0x3456...5d8c
3h ago
Out
13,377 SOL
🔵
0xb6ee...4d95
30m ago
Stake
27,648 SOL

💡 Smart Money

0xb601...075c
Top DeFi Miner
+$3.0M
86%
0xe5eb...9d63
Arbitrage Bot
+$1.7M
78%
0x5e90...420f
Early Investor
-$1.5M
85%