The $40M Illusion: Why Vals AI's Third-Party Evaluation Exposes the Centralization Crisis in AI Trust

AnsemWhale Law

I used to think that the rise of third-party evaluation was the first step toward a more honest AI ecosystem. Then I read the fine print on Vals AI's $40 million Series A, led by a16z, at a $400 million valuation. The numbers are impressive—until you realize that the entire narrative rests on a single, unverifiable claim: model cards from OpenAI, Anthropic, and Google supposedly cite Vals's results. The charts won't tell you this, but the real story is not about the funding. It's about a deeper rot in the way we build trust in AI, and why the blockchain community should be paying attention.

Here is what the numbers won't show: the income growth of "8x" is a self-reported figure with ambiguous wording. The article states that Vals's revenue this year has reached eight times its 2025 annual revenue—a statement that, at face value, contains a temporal contradiction, likely meaning it's eight times the initial projection for 2025. That is not a data point; it's a marketing gloss. The confidence in the entire analysis, as the source report itself admits, is a C. This is not journalism. This is a funded press release dressed as a news item.

Context: The Problem We Need to Solve

The AI evaluation landscape is broken. Public benchmarks like GSM8K and HumanEval have been contaminated—model training sets routinely include test data, and developers optimize directly for the leaderboard. This is not a secret. It is the dirty open secret of the industry. The response has been the rise of "dynamic evaluation" platforms like Vals, which claim to build private, customized test sets from real-world tasks. Their approach is clever: extract historical pull requests from any GitHub repository, turn them into hidden tests, and evaluate a model's ability to complete the task without the model ever seeing the answer.

On the surface, this is a welcome shift. It moves evaluation from static academic benchmarks to the messy, real-world workflow of software development. Vals also extends this to finance, law, and medicine, aiming to verify a model's "production readiness" across domains. This is exactly the kind of infrastructure an industry needs when the trust in self-reported performance has collapsed.

But here is the part the charts don't show: the entire evaluation is still controlled by a single, centralized entity. Vals owns the test generation, the grading, and the results. There is no on-chain verification, no open-source audit of the evaluation pipeline, no decentralized consensus on whether a model actually passed a test. The model vendors—OpenAI, Anthropic, Google—are both the clients and the subjects of evaluation. If Vals charges these companies for evaluation services, the conflict of interest is profound. The auditor becomes beholden to the audited.

Core: The Technical Analysis and the Values Gap

Let me walk you through the technical architecture from a code-integrity perspective. Vals's core innovation is not a new algorithm; it's a productization of dynamic evaluation, similar to the SWE-bench approach. The engineer selects a GitHub repository, pulls historical PRs, and uses the hidden test to check if the model's output matches the developer's intent. This is engineering-level innovation, not research-level. It is valuable, but it is not a moat.

The real question is whether the evaluation results are verifiable. The article notes that Vals does not disclose its task generation methodology, its anti-contamination mechanisms, or whether the historical PRs come from public repositories that may overlap with training data. The risk of data contamination is not eliminated; it is merely deferred. If the model has seen the same PRs during training, the hidden test is not hidden at all.

Based on my own experience auditing smart contracts in 2017, I learned that the difference between a vulnerable system and a secure one is not the sophistication of the code, but the transparency of the audit trail. I submitted 12 critical logic flaws in Gnosis Safe's multi-signature implementation because I could trace every line of code back to its functional intent. Vals's evaluation is a black box. There is no way for an external observer to confirm that the test set is truly private, that the grading is fair, or that the results haven't been manipulated by the very entities being evaluated.

This is where the blockchain ethos becomes essential. The technology we have built—verifiable computation, zero-knowledge proofs, on-chain evidence—can solve precisely this problem. Imagine an evaluation protocol where the test set is committed to a smart contract, the model's output is submitted on-chain, and the result is computed by a decentralized network of validators. The evaluation becomes transparent, immutable, and immune to tampering by any single party. Vals is a step in the right direction, but it is a step toward a centralized solution to a decentralized problem.

Contrarian: The Pragmatism Test

Now, the contrarian angle. I am a believer in slow tech, in building systems that are resilient and ethical. But I also understand the pragmatism of the market. Vals is addressing a real pain point: the cost of evaluating models for enterprise customers. Its model of "plug in your own GitHub repo, get a score" lowers the information asymmetry for buyers. The fact that a16z is putting $40 million into this shows that the venture capital ecosystem sees evaluation as the next infrastructure layer. That is not wrong.

What is wrong is the silence on the centralization of trust. The article's analysis highlights that the "model card citations" are unverified, that the revenue growth is ambiguous, and that the company may be serving a handful of large clients rather than a broad base. The confidence level of C is not a coincidence. It reflects the fact that we are being asked to trust a single company to be the arbiter of AI quality, without any independent oversight.

In the blockchain world, we have learned the hard way that trust in a single entity is fragile. The 2022 collapse of Terra-Luna taught me that when the market turns, centralized structures crumble. The same risk applies here. If Vals's evaluation is shown to be flawed, the entire model card ecosystem built on top of it will collapse. The industry will have wasted years of adoption on a foundation of sand.

Takeaway: The Vision Forward

If you can, I encourage you to look beyond the funding announcement and ask the deeper question: Who guards the guardians? The AI evaluation industry is at a fork. One path leads to a handful of centralized evaluation companies, controlled by the same venture capital that funds the model vendors. The other path leads to a decentralized, verifiable evaluation protocol that anyone can audit, anyone can contribute to, and anyone can trust without asking for permission.

I am not against Vals AI. I am against the illusion that a centralized third-party is the final answer. The blockchain community has the tools to build a better way. The question is whether we have the will to use them.

Follow the fear, not the chart. The fear here is that we are building a new centralized gatekeeper in the name of decentralization. The chart shows a $400 million valuation. The fear shows a system that still trusts a single point of failure. The next wave of AI infrastructure must be built on verifiable, on-chain trust. Otherwise, we are just repeating the same mistakes, but with better models.

Market Prices

BTC Bitcoin
$76,647.4 -1.57%
ETH Ethereum
$2,372.37 -3.17%
SOL Solana
$98.87 -3.21%
BNB BNB Chain
$683.5 -0.34%
XRP XRP Ledger
$1.33 -2.88%
DOGE Dogecoin
$0.0808 -1.83%
ADA Cardano
$0.1947 -1.17%
AVAX Avalanche
$7.12 -1.43%
DOT Polkadot
$0.8532 -0.19%
LINK Chainlink
$11.04 -2.62%

Fear & Greed

63

Greed

Market Sentiment

Event Calendar

{{年份}}
22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

12
05
halving BCH Halving

Block reward halving event

18
03
unlock Sui Token Unlock

Team and early investor shares released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

28
03
unlock Arbitrum Token Unlock

92 million ARB released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

Market Cap

All →
1
Bitcoin
BTC
$76,647.4
1
Ethereum
ETH
$2,372.37
1
Solana
SOL
$98.87
1
BNB Chain
BNB
$683.5
1
XRP Ledger
XRP
$1.33
1
Dogecoin
DOGE
$0.0808
1
Cardano
ADA
$0.1947
1
Avalanche
AVAX
$7.12
1
Polkadot
DOT
$0.8532
1
Chainlink
LINK
$11.04

Tools

All →

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

🐋 Whale Tracker

🟢
0x1e31...1c05
6h ago
In
1,708.21 BTC
🟢
0x9239...5e71
1d ago
In
3,278,035 DOGE
🔴
0x92b9...b4e0
12m ago
Out
20,663 SOL

💡 Smart Money

0xf997...dc76
Arbitrage Bot
+$1.9M
92%
0xf528...d64d
Arbitrage Bot
+$0.1M
80%
0x124b...069a
Market Maker
+$1.1M
83%