Inference Optimization
Search documents
X @Avi Chawla
Avi Chawla· 2026-08-13 19:57
Continuous batching in LLMs, clearly explained:(a popular LLM interview question; bookmark this)In traditional ML inference, a batch is a matrix.Every input is padded to the same length, one forward pass runs, and every row finishes at the same moment.LLM decoding does not work that way.One forward pass produces one token per sequence, so a request needs as many passes as it has output tokens, and nobody knows that count until the model emits a stop token.Under static batching, membership is fixed when the ...
We're in a ‘Golden Age of Technology,’ Says IVP's Wilhelm
Bloomberg Technology· 2026-08-13 17:54
I wanna start with general question about buying and building organically. You know, at all layers of the five layer cake stack, all of that's happening at once. Some people are building. Some people are are out in the market.What is it that you see most in your in your portfolio. Yeah. Good question.So we're a venture capitalist. Right. We're trying to invest in all these companies.We're you know, some of them we're gonna sell to bigger companies, but the ones that we love are the ones that go all the way. ...