vLLM 0.27 Release: New CondensePyramid V1 Attention Kernel and Multi-Modal Token Limits
vLLM 0.27 adds CondensePyramid V1 attention for multi-modal long context, new token-tier limits on templates, faster ASR CPU preprocessing via multi-threading, and CPU W4A16 INT4 MoE support.