
GPU compiler and cross-cloud compute platform for AI
sf-tensor.com (opens in a new tab)SF Tensor rebuilds the AI compute stack from the chip up to make training faster, cheaper and less tied to any one vendor. Its Kernel Optimizer rewrites code into the fastest form for a given chip and cluster layout, and its Model Foundry manages training runs and shifts workloads between clouds and chips as price and availability change. The compiler checks correctness at the end of optimisation rather than at each step, which lets it search far more options, and it tops NVIDIA's own benchmark, which spans hundreds of real production kernels. A three-person team from the company once ran foundation model pre-training across 4,000 AMD GPUs, and its kernels have run AlphaFold 3 pre-training at 3.4 times the usual throughput. Most work happens in its San Francisco office.