T
traeai
Sign in

人物

AK

别名:@_akhaliq

推文作者,分享了相关论文信息。

已跟踪 30 条高相关材料

TraeAI 观察

相关材料

已收录 30 条与 AK 相关的内容,按评分排序。

SkillOS

Learning Skill Curation for Self-Evolving Agents

paper: https://t.co/C6yKe6Kuou

SkillOS: Learning Skill Curation for Self-Evolving Agents

AK(@_akhaliq)60 字 (约 1 分钟)
78

SkillOS is a skill orchestration system for self-evolving agents, achieving 34% higher accuracy through dynamic skill library and meta-learning mechanisms.

入选理由:SkillOS 采用动态技能库,支持实时技能增删与更新。

FeaturedTweet#AI Agent#Skill Curation#Self-Evolving Systems#Meta-Learning英文
GPU Forecasters

Language Models as Selective Surrogates for Kernel Runtime Optimization

This article explores a new approach to GPU kernel runtime optimization using language models as selective surrogates, achieving significant performance improvements by predicting and selecting optimal kernel configurations.

入选理由:语言模型被用作选择性代理,预测 GPU 内核的最佳配置。

FeaturedTweet#GPU#Language Models#Kernel Optimization#Runtime Performance#AI Acceleration英文
Seeing Isn't Knowing

Do VLMs Know When Not to Answer Spatial Questions (and Why)?

Seeing Isn't Knowing: The Limitations of VLMs in Spatial Reasoning

AK(@_akhaliq)53 字 (约 1 分钟)
75

This article explores the limitations of Visual Language Models (VLMs) in handling spatial questions, highlighting their tendency to confidently generate answers even when visual cues are ambiguous, and suggests introducing uncertainty mechanisms to improve model robustness.

入选理由:VLMs 在缺乏明确视觉线索时,仍可能自信地生成空间问题的答案。

FeaturedTweet#VLM#Visual Language Model#Spatial Reasoning#Uncertainty#AI Explainability英文
LongMINT

Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

LongMINT

AK(@_akhaliq)57 字 (约 1 分钟)
75

LongMINT is a new benchmark testing framework for evaluating memory capabilities under multi-target interference in long-horizon agent systems, which has gained attention through academic sharing on Twitter. This framework specifically addresses memory interference issues in AI agents during long-term tasks and provides standardized testing methods for measuring continuous learning and memory management capabilities of agent systems.

入选理由:LongMINT是专门评估长视界智能体记忆干扰的新基准测试框架

FeaturedTweet#LongMINT#AI Agents#Memory Evaluation#Benchmarking英文
Mix-Quant

Quantized Prefilling, Precise Decoding for Agentic LLMs

Mix-Quant

AK(@_akhaliq)44 字 (约 1 分钟)
75

Mix-Quant technology significantly improves the efficiency and precision balance of agentic LLMs through a hybrid strategy of quantized prefilling and precise decoding, providing new optimization directions for large model deployment.

入选理由:Mix-Quant采用量化预填充和精确解码的混合策略优化LLM性能

FeaturedTweet#Mix-Quant#LLM#Quantization Technology#AI Inference英文
MulTaBench

Benchmarking Multimodal Tabular Learning with Text and Image

MulTaBench

AK(@_akhaliq)54 字 (约 1 分钟)
75

MulTaBench is a benchmark for evaluating multimodal tabular learning with text and image.

入选理由:MulTaBench 包含 12 个数据集和 3 种任务类型。

FeaturedTweet#Multimodal Learning#Tabular Data中文
paper: https://t.co/eG6d4D9oEF

AK on X: "paper: https://t.co/eG6d4D9oEF"

AK(@_akhaliq)45 字 (约 1 分钟)
75

The article introduces the MACE-Dance model for music-driven dance video generation.

入选理由:MACE-Dance 是一种音乐驱动的舞蹈视频生成模型。

FeaturedTweet#AI#video generation#music-driven#deep learning中文
ESI-Bench

Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

ESI-Bench is a novel benchmark focused on evaluating embodied spatial intelligence models in perception-action loops, offering more challenging scenarios and metrics than existing tests.

入选理由:ESI-Bench 采用连续 3D 轨迹预测任务,比现有基准更具挑战性

FeaturedTweet#Embodied Intelligence#Spatial Intelligence#AI Benchmark#3D Trajectory Prediction#Perception-Action Loop英文
Do Enterprise Systems Need Learned World Models? 

The Importance of Context to Infer Dynamics

企业系统是否需要学习世界模型?文章探讨了上下文对推断动态的重要性,强调了在复杂环境中理解背景信息的价值。

入选理由:在企业系统中,上下文对于推断系统的动态行为至关重要。

FeaturedTweet#企业系统#世界模型#上下文#动态推断英文
PhyMotion

Structured 3D Motion Reward for Physics-Grounded Human Video Generation

PhyMotion

AK(@_akhaliq)42 字 (约 1 分钟)
65

PhyMotion introduces a structured 3D motion reward mechanism grounded in physics to enhance the realism of human video generation.

入选理由:PhyMotion 引入物理约束以增强视频生成的真实性。

FeaturedTweet#AI#Video Generation英文
VideoChat3

Fully Open Video MLLM for Efficient and Generalist Video Understanding

VideoChat3是首个全开放的视频多模态大模型,支持高效通用视频理解,但技术细节披露有限。

入选理由:VideoChat3是首个全开放的视频MLLM,支持高效视频理解

FeaturedTweet#VideoChat3#MLLM#视频理解#开源英文
ViQ

Text-Aligned Visual Quantized Representations at Any Resolution

ViQ Text-Aligned Visual Quantized Representations at Any Resolution

AK(@_akhaliq)54 字 (约 1 分钟)
60

ViQ 是一种文本对齐的视觉量化表示方法,可在任意分辨率下使用。

入选理由:ViQ 支持任意分辨率的视觉量化表示。

FeaturedTweet#ViQ#视觉量化#文本对齐#AI英文
Confidence-Aware Tool Orchestration for Robust Video Understanding

Confidence-Aware Tool Orchestration for Robust Video Understanding

AK(@_akhaliq)51 字 (约 1 分钟)
60

本文提出了一种基于置信度的工具编排方法,用于提升视频理解的鲁棒性,但内容较为简略,缺乏具体实现细节。

入选理由:置信度感知的工具编排方法可提升视频理解的鲁棒性。

FeaturedTweet#视频理解#工具编排#AI英文
LoopCoder-v2

Only Loop Once for Efficient Test-Time Computation Scaling

LoopCoder-v2 Only Loop Once for Efficient Test-Time Computation Scaling

AK(@_akhaliq)67 字 (约 1 分钟)
60

LoopCoder-v2 是一种优化测试时计算效率的方法,通过减少循环次数提升性能。

入选理由:LoopCoder-v2 通过减少循环次数来提高测试时的计算效率。

FeaturedTweet#LoopCoder-v2#计算优化#测试效率英文
paper: https://t.co/NluxzaDkCS

paper: https://t.co/NluxzaDkCS

AK(@_akhaliq)43 字 (约 1 分钟)
60

文章分享了一篇关于LoopCoder-v2的论文,旨在提高测试时计算效率。

入选理由:LoopCoder-v2通过仅循环一次来提高测试时计算效率。

FeaturedTweet#LoopCoder-v2#Hugging Face#AI#论文中英混合
World Tracing

Generative Pixel-Aligned Geometry Beyond the Visible

World Tracing Generative Pixel-Aligned Geometry Beyond the Visible

AK(@_akhaliq)65 字 (约 1 分钟)
60

World Tracing 是一种生成像素对齐几何的新技术,但文章内容信息密度低,缺乏具体机制和实用价值。

入选理由:World Tracing 是一种生成像素对齐几何的新技术。

FeaturedTweet#AI#计算机视觉#生成模型英文
μ_0

A Scalable 3D Interaction-Trace World Model

μ_0 A Scalable 3D Interaction-Trace World Model

AK(@_akhaliq)62 字 (约 1 分钟)
60

文章介绍了一种可扩展的3D交互轨迹世界模型μ_0,但内容信息密度低,缺乏具体技术细节和实用价值。

入选理由:文章提出了一种名为μ_0的3D交互轨迹世界模型。

FeaturedTweet#3D模型#AI#世界模型英文
CHORUS

Decentralized Multi-Embodiment Collaboration with One VLA Policy

CHORUS Decentralized Multi-Embodiment Collaboration with One VLA Policy

AK(@_akhaliq)65 字 (约 1 分钟)
60

CHORUS 是一种基于单一 VLA 策略的去中心化多实体协作方法,但文章内容信息密度低,缺乏具体机制和实践细节。

入选理由:CHORUS 采用单一 VLA 策略实现多实体协作。

FeaturedTweet#AI#协作#VLA#去中心化英文
paper: https://t.co/aID0K3TdFx

paper: https://t.co/aID0K3TdFx

AK(@_akhaliq)45 字 (约 1 分钟)
60

文章分享了一篇关于SpenseGPT的论文,探讨了一种名为SpenseGPT的模型,旨在通过稀疏和密集GEMMs实现大语言模型的高效推理。

入选理由:SpenseGPT是一种通过稀疏和密集GEMMs实现高效推理的模型。

FeaturedTweet#SpenseGPT#LLM#GEMMs#Hugging Face英文
On the Geometry of On-Policy Distillation

On the Geometry of On-Policy Distillation

AK(@_akhaliq)51 字 (约 1 分钟)
60

文章探讨了On-Policy Distillation的几何特性,但信息密度较低,缺乏具体实践指导。

入选理由:文章讨论了On-Policy Distillation的几何特性。

FeaturedTweet#On-Policy Distillation#机器学习#几何特性英文
Latent Spatial Memory for Video World Models

Latent Spatial Memory for Video World Models

AK(@_akhaliq)55 字 (约 1 分钟)
60

文章介绍了一种用于视频世界模型的潜在空间记忆方法,但信息密度较低,缺乏具体机制和实践指导。

入选理由:潜在空间记忆方法被提出用于视频世界模型。

FeaturedTweet#视频世界模型#潜在空间记忆#AI研究英文
paper: https://t.co/5HPrTxJGK9

paper: https://t.co/5HPrTxJGK9

AK(@_akhaliq)46 字 (约 1 分钟)
60

文章推荐了一篇关于企业系统是否需要学习世界模型的研究论文,探讨了上下文对推理的重要性。

入选理由:论文《Do Enterprise Systems Need Learned World Models?》探讨了企业系统中学习世界模型的需求。

FeaturedTweet#企业系统#世界模型#上下文推理#AI研究#论文英文
Explorative Modeling

Unlocking a Third Pretraining Axis and End-to-End Generation

paper: https://t...

推文分享了一篇关于探索性建模的新论文,提出预训练模型的第三维度和端到端生成方法,但未提供技术细节。

入选理由:推文分享了一篇关于探索性建模的新论文,提出预训练模型的第三维度和端到端生成方法,但未提供技术细节

FeaturedTweet#预训练模型#生成模型#AI研究中英混合

跨材料问答 · AK

回答基于:AK 相关 30 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.