1:23:36Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Neil Movva, co-founder of Sail Research and a former NVIDIA GPU kernel engineer, tells Patrick O'Shaughnessy how he intends to drive the cost of a token down by three to six orders of magnitude. His argument runs the whole stack: rebuild inference around throughput instead of latency, buy any chip at the right price rather than bidding against the frontier labs for Blackwell, and scavenge power by accepting distributed one megawatt data centers that only stay up 95 percent of the time. Along the way he gives an unusually clear technical account of tensor cores, SRAM versus HBM, the KV cache, why attention is memory bound while the MLP is compute bound, and why performance per watt has barely moved from Hopper to Blackwell to Rubin. The bet only works because background agents run for hours and nobody is waiting on the answer.