1 writing found
Speculative decoding for vision-language models achieves significant inference speedups on edge devices and GPUs with minimal parameter overhead and day-one framework support.