
Optimised inference for open-weight models
wafer.ai (opens in a new tab)Wafer runs dedicated and serverless inference for open-weights models, using AI to optimise AI infrastructure — tuning the entire serving stack so the same model runs faster on cheaper hardware, with customers keeping the saving. That last detail is the business model rather than a nicety: the company's value is measured directly in the difference between what inference costs and what it should cost. Wafer is seven people in San Francisco and intends to double as quickly as it can hire, selling into AI-native companies running production inference.
