T
traeai
Sign in

概念

NVFP4

别名:NV FP4

NVIDIA推出的混合精度计算格式

已跟踪 9 条高相关材料

TraeAI 观察

相关材料

已收录 9 条与 NVFP4 相关的内容,按评分排序。

What Is NVFP4? Faster LLM Inference Without Losing Quality

What Is NVFP4? Faster LLM Inference Without Losing Quality

NVIDIA Developer1253 字 (约 6 分钟)
85

NVFP4通过4位浮点数格式和双缩放策略,在减少内存占用的同时保持模型质量,显著提升大语言模型推理效率。

入选理由:NVFP4使用E2M1格式,每个权重仅需4位,内存减少75%

FeaturedVideo#大语言模型#量化技术#NVIDIA#FP4#模型优化英文
The Secret to Cheaper LLM Inference: NVFP4

The Secret to Cheaper LLM Inference: NVFP4

NVIDIA Developer186 字 (约 1 分钟)
85

NVFP4通过4位量化技术降低大模型内存占用,在保持质量前提下提升推理效率,使更大模型或更高并发成为可能。

入选理由:NVFP4量化格式可减少模型内存占用达50%以上

FeaturedVideo#LLM#推理优化#量化技术#NVIDIA英文
We evaluated Nemotron-3-Embed from @NVIDIAAI for our @mem0ai memory retrieval pipeline

Tested on lo...

mem0团队采用Nemotron-3-Embed模型后,长文本检索准确率提升1.67个百分点,其开放权重和NVFP4加速特性成为关键选择因素。

入选理由:Nemotron-3-Embed在longmemeval测试中实现80.38的检索准确率,超越qwen-3-600m模型

FeaturedTweet#NVIDIA#Nemotron-3-Embed#mem0ai#机器学习#嵌入模型中英混合
Holo3.1: Fast & Local Computer Use Agents

Holo3.1: Fast & Local Computer Use Agents

Hugging Face Blog808 字 (约 4 分钟)
85

Holo3.1 is Hugging Face's new computer-use agent model supporting cross-platform, multi-framework deployment and first releasing quantized weights (FP8/Q4 GGUF/NVFP4) for local inference.

入选理由:Holo3.1 在 AndroidWorld 上 35B-A3B 模型准确率从 67% 提升至 79.3%

FeaturedArticle#Computer Use Agent#Hugging Face#Quantized Model#Mobile Automation英文
NVIDIA Nemotron 3 Ultra now available on Amazon SageMaker JumpStart

NVIDIA Nemotron 3 Ultra Now Available on Amazon SageMaker JumpStart

AWS Machine Learning Blog952 字 (约 4 分钟)
82

NVIDIA Nemotron 3 Ultra is now available on Amazon SageMaker JumpStart with one-click deployment. This 550B-parameter MoE model is designed for long-running agents, delivering 5x faster inference, 30% lower cost, and 1M token context support.

入选理由:Nemotron 3 Ultra采用混合Transformer-Mamba MoE架构,550B总参仅激活55B,显著降低Agent任务计算开销。

FeaturedArticle#Nemotron 3 Ultra#SageMaker JumpStart#Agentic AI#MoE#AWS英文
Long video generation is a systems problem.

Introducing LongLive-2.0 from NVIDIA Research: an end-t...

NVIDIA Research releases LongLive-2.0 system that adopts end-to-end NVFP4 training and inference architecture to solve long video generation problems, eliminating model deployment gaps through unified training-inference precision alignment while improving speed and memory efficiency.

入选理由:LongLive-2.0采用NVFP4低精度训练推理架构

FeaturedTweet#NVIDIA#Video Generation#Low Precision Computing#AI Systems英文
who's working on an NVFP4 version of Kimi-K3?

who's working on an NVFP4 version of Kimi-K3?

Julien Chaumond(@julien_c)177 字 (约 1 分钟)
60

关于Kimi-K3的NVFP4版本开发存在争议,部分开发者认为其意义有限,但Red Hat已推出相关优化方案。

入选理由:Eric认为NVFP4版本对数据中心外用户无实质价值

FeaturedTweet#Kimi-K3#NVFP4#Blackwell GPU#Red Hat AI中英混合
Nvidia presents LongLive-2.0

An NVFP4 Parallel Infrastructure for Long Video Generation

Nvidia presents LongLive-2.0

AK(@_akhaliq)52 字 (约 1 分钟)
45

Nvidia releases LongLive-2.0, an NVFP4 parallel infrastructure for long video generation, but the tweet only announces the product name without disclosing any technical implementation details.

入选理由:Nvidia发布LongLive-2.0长视频生成基础设施

FeaturedTweet#Nvidia#Video Generation#NVFP4#Parallel Computing#AI Infrastructure英文

跨材料问答 · NVFP4

回答基于:NVFP4 相关 9 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.