T
traeai
Sign in

概念

Speculative Decoding

别名:假设采样

基于候选序列验证的并行推理技术

已跟踪 4 条高相关材料

TraeAI 观察

相关材料

已收录 4 条与 Speculative Decoding 相关的内容,按评分排序。

NVIDIA AI(@NVIDIAAI) 图标

Read the technical blog for the five guidelines: https://t.co/HaXwdiTYYB

NVIDIA AI(@NVIDIAAI)54 字 (约 1 分钟)
85

NVIDIA 技术博客提出通过 Speculative Decoding 优化大模型推理的五项指南,可使推理速度提升 2-5 倍。

入选理由:Speculative Decoding 技术可使 LLM 推理速度提升 2-5 倍

FeaturedTweet#LLM#Speculative Decoding#NVIDIA#AI模型优化英文
Apple Machine Learning Research 图标

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

Apple Machine Learning Research493 字 (约 2 分钟)
85

ARBITRAGE通过动态路由机制提升推理效率,在保持精度前提下减少推理延迟达2倍。

入选理由:ARBITRAGE框架使推理延迟降低约2倍,精度保持不变

FeaturedArticle#Large Language Models#Speculative Decoding#Efficient Reasoning#Machine Learning Research#Apple英文
I've seen some confusion online on how to run llama.cpp with MTP (Multi-token prediction) in the sim...

How to Run llama.cpp with MTP (Multi-token Prediction)

Julien Chaumond(@julien_c)255 字 (约 2 分钟)
75

MTP is a new speculative decoding feature built into llama.cpp that can approximately double token generation speed for most use cases, achieving ~30 tok/sec with the Dense 27B model and ~100 tok/sec with the MoE model.

入选理由:MTP是内置于模型本身的投机解码新特性,可将token生成速度提升约2倍

FeaturedTweet#llama.cpp#MTP#Speculative Decoding#Qwen#LLM Inference Optimization英文
RL post-training is hitting a rollout bottleneck. 

This new paper from #NVIDIAResearch shows how sp...

NVIDIA 研究提出将 speculative decoding 引入 NeMo-RL + vLLM 架构,实现 RL 后训练 rollout 阶段无损加速:8B 模型吞吐提升 1.8 倍,235B 模型端到端预计提速 2.5 倍。

入选理由:RLHF/RLAIF 后训练的 rollout 阶段已成为性能瓶颈

FeaturedTweet#RLHF#speculative decoding#vLLM#NeMo-RL#NVIDIA中英混合

跨材料问答 · Speculative Decoding

回答基于:Speculative Decoding 相关 4 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.