
Enterprise AI model evaluation
vals.ai (opens in a new tab)Vals AI evaluates AI models for enterprises. Its product, Vals Smith, runs against a customer's own code repositories and reports which model performs best on that customer's work, giving buyers a defensible basis for model spending that can reach eight figures. Every major foundation model lab works with Vals, as do banks and hospital networks among the biggest anywhere. Its benchmark studies reach the press, including a finding that AI tools mostly fail at basic financial tasks and a ranking of vibe-coding tools. The small team has worked at NVIDIA, Meta, Microsoft and Palantir, and its members' published research has drawn more than 300 citations between them. Early hires include Stanford PhDs, former Jane Street quants and Snorkel's first designer.