When a16z writes a $40M check for an 'independent' AI evaluator, the code doesn't lie — but the narrative might. Vals AI just closed a Series A at a $400M valuation, promising to fix the broken benchmark economy by pulling real tasks from GitHub pull requests. The premise is seductive: model evaluation that actually reflects production work, not memorized exam questions. But look closer, and the arbitrage between marketing and technical reality reveals itself. The code doesn't lie, but the claims around it do.
Context: Why Now?
The AI evaluation market is a mess. Public benchmarks like GSM8K and HumanEval are contaminated — training data leaks are common, and model vendors optimize for leaderboard scores rather than real-world utility. Enterprise buyers are desperate for a signal that a model can actually handle their codebase, their legal documents, their financial data. Into this vacuum steps Vals AI, founded by ex-Google engineers, offering a platform that extracts tasks from a company's own GitHub history and tests models against hidden tests. It's SWE-bench as a service, with a SaaS wrapper.
But here's the first red flag: the entire evaluation pipeline is proprietary. The task generation, the hidden test creation, the scoring mechanism — all black box. The company says it can evaluate across domains like finance, law, and medicine. Yet no methodology paper has been published. No independent audit of the evaluation set has been conducted. The company claims OpenAI, Anthropic, Google, Meta, and xAI cite its results in their model cards. But that's a self-reported claim, unverifiable without subpoena-level access.

Core: The Numbers and the Noise
Let's parse the hard data. $40M at $400M post-money implies roughly 9-10% dilution for the Series A — standard for a growth-stage round. But the valuation is rich for a company whose revenue is not only undisclosed but described with a temporal paradox. The statement "this year's revenue has reached 8 times the 2025 full-year revenue" is physically impossible if we are in 2025. The most charitable interpretation: "revenue in the current period (likely a partial year) is 8x the company's initial internal forecast for the entire calendar year 2025." That's a different, weaker claim. It could also mean year-over-year growth of 8x, but that would be typical for a startup from a small base. The ambiguity is a deliberate fog machine.
We didn't expect the financial story to be this opaque. The revenue figure is not in ARR, not in contract value, not in billings. It's a single unverifiable number with a question mark over its definition. The company also hasn't disclosed customer count, average contract size, retention rates, or net dollar expansion. For a company valued at $400M, that's a data desert. The models are getting evaluated, but the company itself isn't.
On the technical side, Vals's approach is clever but not new. SWE-bench introduced dynamic evaluation from real GitHub issues. Vals is productizing that concept. But the core innovation is operational, not algorithmic. The real moat — if there is one — would be the proprietary dataset of hidden tasks and the trust of enterprise clients. But trust is fragile when the evaluator is funded by a16z, which also funds many of the AI companies being evaluated. Smart contracts are smart; humans are the bug. The conflict of interest is obvious: can an evaluator remain independent when its largest investor also backs the companies it judges? The article doesn't even mention this.
Contrarian: The Unreported Angle
Here's what the mainstream coverage misses: Vals AI's evaluation method may actually be less robust than it claims. The company extracts tasks from "historical pull requests" from any GitHub repository. But if the repository is public, those PRs are part of the training data for many large language models. The models may have seen the solution code during training. Vals claims it uses "hidden tests" to avoid this, but if the test setup is derived from the same PR, the model could still pattern-match. The only way to prevent contamination is to use private, proprietary codebases that the model has never seen. That's feasible for enterprise clients, but the company's public demo material uses open-source repos. The risk of evaluation contamination is high, and the company's silence on this is deafening.
Arbitrage is just patience wearing a speed suit. The real arbitrage here is information asymmetry: Vals knows its own methodology, but the market is buying the story without technical verification. The hidden information is that Vals likely needs heavy human annotation for legal and medical tasks — the cost structure is not pure software. The company hasn't disclosed how many evaluators it employs or how it scales quality control. For a $400M company, that's a material risk.
Another blind spot: the model card citations. If Vals is indeed referenced by OpenAI and others, those citations are likely part of a broader evaluation suite, not an endorsement of Vals as the sole arbiter. The company may be overstating its influence. Moreover, the regulatory angle: in regulated industries like finance and healthcare, can a third-party evaluation from a VC-backed startup be used as compliance evidence? Unclear. The article's source is a blockchain monitoring channel, suggesting the crypto crowd is paying attention — but this is a pure AI story. The industry impact is real, but the blockchain angle is tenuous at best.
Takeaway: Who Evaluates the Evaluator?
Vals AI is a bet on the commoditization of trust. But trust is not a commodity; it's a relationship. The company's success depends on maintaining independence, methodology transparency, and rigorous data hygiene. The funding round is a signal that the market wants a third-party evaluation standard. But the current state of transparency is abysmal. Floor prices are opinions; volume is the truth. Without open-sourced evaluation sets, published audits, or revenue verification, the $400M valuation is an opinion, not a fact. The next 12 months will reveal whether Vals can build the infrastructure to match the narrative. The code doesn't lie — but the company's claims need to be put to the test.

For now, the smart money is waiting. Liquidity leaves fast, but the smart money stays.