btw this is the same trick that makes spec decoding work
if we have a small model propose the next n tokens, we can check them in one forward pass! as long as the small model is calibrated (usually in agreement) with the larger one, you should see an increase in inference speed because you save on expensive large-model forward passes!
incredibly clever chunking trick by the zeroentropy team (now at notion)
llms are autoregressive, which means generating text is super slow.
but passing text in and finding the logprobs of a particular token at all locations takes a single forward pass!
so we can ask the llm, 'like yo llm, please chunk this by putting this unlikely character 段 wherever there should be chunks"
and we instantly get the segmentation locations (well, not instant, we still need a forward pass which means we need to prefill)
but the advantage is super accurate chunks in the time it takes to prefill!