Ethereum

Vals AI's $40M Raise: The Illusion of Independent AI Evaluation

CryptoNode

When a16z writes a $40M check for an 'independent' AI evaluator, the code doesn't lie — but the narrative might. Vals AI just closed a Series A at a $400M valuation, promising to fix the broken benchmark economy by pulling real tasks from GitHub pull requests. The premise is seductive: model evaluation that actually reflects production work, not memorized exam questions. But look closer, and the arbitrage between marketing and technical reality reveals itself. The code doesn't lie, but the claims around it do.

Context: Why Now?

The AI evaluation market is a mess. Public benchmarks like GSM8K and HumanEval are contaminated — training data leaks are common, and model vendors optimize for leaderboard scores rather than real-world utility. Enterprise buyers are desperate for a signal that a model can actually handle their codebase, their legal documents, their financial data. Into this vacuum steps Vals AI, founded by ex-Google engineers, offering a platform that extracts tasks from a company's own GitHub history and tests models against hidden tests. It's SWE-bench as a service, with a SaaS wrapper.

But here's the first red flag: the entire evaluation pipeline is proprietary. The task generation, the hidden test creation, the scoring mechanism — all black box. The company says it can evaluate across domains like finance, law, and medicine. Yet no methodology paper has been published. No independent audit of the evaluation set has been conducted. The company claims OpenAI, Anthropic, Google, Meta, and xAI cite its results in their model cards. But that's a self-reported claim, unverifiable without subpoena-level access.

Vals AI's $40M Raise: The Illusion of Independent AI Evaluation

Core: The Numbers and the Noise

Let's parse the hard data. $40M at $400M post-money implies roughly 9-10% dilution for the Series A — standard for a growth-stage round. But the valuation is rich for a company whose revenue is not only undisclosed but described with a temporal paradox. The statement "this year's revenue has reached 8 times the 2025 full-year revenue" is physically impossible if we are in 2025. The most charitable interpretation: "revenue in the current period (likely a partial year) is 8x the company's initial internal forecast for the entire calendar year 2025." That's a different, weaker claim. It could also mean year-over-year growth of 8x, but that would be typical for a startup from a small base. The ambiguity is a deliberate fog machine.

We didn't expect the financial story to be this opaque. The revenue figure is not in ARR, not in contract value, not in billings. It's a single unverifiable number with a question mark over its definition. The company also hasn't disclosed customer count, average contract size, retention rates, or net dollar expansion. For a company valued at $400M, that's a data desert. The models are getting evaluated, but the company itself isn't.

On the technical side, Vals's approach is clever but not new. SWE-bench introduced dynamic evaluation from real GitHub issues. Vals is productizing that concept. But the core innovation is operational, not algorithmic. The real moat — if there is one — would be the proprietary dataset of hidden tasks and the trust of enterprise clients. But trust is fragile when the evaluator is funded by a16z, which also funds many of the AI companies being evaluated. Smart contracts are smart; humans are the bug. The conflict of interest is obvious: can an evaluator remain independent when its largest investor also backs the companies it judges? The article doesn't even mention this.

Contrarian: The Unreported Angle

Here's what the mainstream coverage misses: Vals AI's evaluation method may actually be less robust than it claims. The company extracts tasks from "historical pull requests" from any GitHub repository. But if the repository is public, those PRs are part of the training data for many large language models. The models may have seen the solution code during training. Vals claims it uses "hidden tests" to avoid this, but if the test setup is derived from the same PR, the model could still pattern-match. The only way to prevent contamination is to use private, proprietary codebases that the model has never seen. That's feasible for enterprise clients, but the company's public demo material uses open-source repos. The risk of evaluation contamination is high, and the company's silence on this is deafening.

Arbitrage is just patience wearing a speed suit. The real arbitrage here is information asymmetry: Vals knows its own methodology, but the market is buying the story without technical verification. The hidden information is that Vals likely needs heavy human annotation for legal and medical tasks — the cost structure is not pure software. The company hasn't disclosed how many evaluators it employs or how it scales quality control. For a $400M company, that's a material risk.

Another blind spot: the model card citations. If Vals is indeed referenced by OpenAI and others, those citations are likely part of a broader evaluation suite, not an endorsement of Vals as the sole arbiter. The company may be overstating its influence. Moreover, the regulatory angle: in regulated industries like finance and healthcare, can a third-party evaluation from a VC-backed startup be used as compliance evidence? Unclear. The article's source is a blockchain monitoring channel, suggesting the crypto crowd is paying attention — but this is a pure AI story. The industry impact is real, but the blockchain angle is tenuous at best.

Takeaway: Who Evaluates the Evaluator?

Vals AI is a bet on the commoditization of trust. But trust is not a commodity; it's a relationship. The company's success depends on maintaining independence, methodology transparency, and rigorous data hygiene. The funding round is a signal that the market wants a third-party evaluation standard. But the current state of transparency is abysmal. Floor prices are opinions; volume is the truth. Without open-sourced evaluation sets, published audits, or revenue verification, the $400M valuation is an opinion, not a fact. The next 12 months will reveal whether Vals can build the infrastructure to match the narrative. The code doesn't lie — but the company's claims need to be put to the test.

Vals AI's $40M Raise: The Illusion of Independent AI Evaluation

For now, the smart money is waiting. Liquidity leaves fast, but the smart money stays.

Market Prices

BTC Bitcoin
$63,034.9 +0.32%
ETH Ethereum
$1,879.71 +0.25%
SOL Solana
$75.16 -0.87%
BNB BNB Chain
$611.1 +0.63%
XRP XRP Ledger
$1 -0.40%
DOGE Dogecoin
$0.0700 +0.23%
ADA Cardano
$0.1788 -1.97%
AVAX Avalanche
$6.61 +3.23%
DOT Polkadot
$0.7703 +1.64%
LINK Chainlink
$9.3 +6.31%

Fear & Greed

34

Fear

Market Sentiment

7x24h Flash News

More >
{{快讯列表(10)}} {{loop}}
{{快讯时间}}

{{快讯内容}}

{{快讯标签}}
{{/loop}} {{/快讯列表}}

Event Calendar

{{年份}}
28
03
unlock Arbitrum Token Unlock

92 million ARB released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

12
05
halving BCH Halving

Block reward halving event

18
03
unlock Sui Token Unlock

Team and early investor shares released

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

Tools

All →

Altseason Index

44

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$63,034.9
1
Ethereum
ETH
$1,879.71
1
Solana
SOL
$75.16
1
BNB Chain
BNB
$611.1
1
XRP Ledger
XRP
$1
1
Dogecoin
DOGE
$0.0700
1
Cardano
ADA
$0.1788
1
Avalanche
AVAX
$6.61
1
Polkadot
DOT
$0.7703
1
Chainlink
LINK
$9.3

🐋 Whale Tracker

🟢
0x08e3...98a3
12m ago
In
694,176 USDT
🔴
0xa7c7...5ebc
3h ago
Out
48,771 BNB
🔴
0xba41...9efe
30m ago
Out
28,092 BNB

💡 Smart Money

0x7836...8501
Market Maker
+$2.2M
67%
0x1eb0...f401
Arbitrage Bot
+$0.5M
77%
0xec6d...f87c
Experienced On-chain Trader
+$4.7M
75%