T
traeai
Sign in

公司

Together AI

别名:togetherai

发布ThunderAgent的AI公司

已跟踪 23 条高相关材料

TraeAI 观察

相关材料

已收录 23 条与 Together AI 相关的内容,按评分排序。

Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

Together AI optimized the deployment of MiniMax M3, achieving 81–125% throughput improvements through architectural and engineering innovations.

入选理由:MiniMax M3 supports 1M-token context and native multimodality, making it suitable for complex real-world tasks.

FeaturedArticle#MiniMax#M3#Sparse Attention#Multimodality#Inference Optimization英文
Together AI Blog 图标

ThunderAgent通过优化KV缓存管理,实现单节点吞吐量提升2倍,集群加速2.4倍,解决代理推理中的缓存抖动问题。

入选理由:单节点吞吐量提升2.5倍,P50延迟降低10倍

FeaturedArticle#ThunderAgent#合成数据生成#LLM推理优化#KV缓存管理英文
Together AI Blog 图标

Configuring Dedicated Model Inference

Together AI Blog1796 字 (约 8 分钟)
85

Together AI平台通过端点/部署/配置三元组与容量感知路由实现模型推理配置,支持零停机变更和流量实验。

入选理由:配置由端点(固定身份)、部署(模型+硬件组合)、配置(运行配方)三部分构成

FeaturedArticle#模型推理#部署配置#流量管理#AI平台英文
Together AI Blog 图标

Together AI与Moonshot AI合作推出2.8T参数开源模型Kimi K3,支持长上下文和视觉任务,开发者可直接通过其平台部署和微调。

入选理由:Kimi K3是当前参数规模最大的开源模型(2.8T),支持1M token上下文窗口。

FeaturedArticle#AI模型#开源#Moonshot AI#Together AI英文
Together AI Blog 图标

Autoscaling endpoints for LLM inference

Together AI Blog2237 字 (约 9 分钟)
85

Together AI平台通过自适应扩展端点,结合LLM推理特有的指标(如TTFT、GPU利用率),实现更高效的资源管理,减少过量和不足配置的成本。

入选理由:选择TTFT和GPU利用率作为指标可有效平衡资源成本

FeaturedArticle#LLM推理#自适应扩展#Together AI#GPU优化英文
Together AI Blog 图标

Kimi K3: The Complete Developer Guide

Together AI Blog3077 字 (约 13 分钟)
85

Kimi K3是首个2.8万亿参数开源模型,支持1M上下文和高效推理,适用于复杂编程与知识工作。

入选理由:Kimi K3参数量达2.8万亿,是首个3万亿参数级开源模型

FeaturedArticle#Kimi K3#开源模型#AI架构#Together AI英文
Together AI Blog 图标

Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

Together AI Blog1176 字 (约 5 分钟)
85

Kimi K3在DeepSWE基准测试中以1/3成本超越Claude Fable 5,多次尝试后性能反超,但后者可靠性更高。

入选理由:Kimi K3单次任务成本仅4.65美元,为Claude Fable 5的1/3

FeaturedArticle#Kimi K3#Claude Fable 5#DeepSWE#模型比较英文
Together AI Blog 图标

What does 99.9% uptime mean for inference?

Together AI Blog1508 字 (约 7 分钟)
85

不同可靠性级别(99%、99.9%、99.99%)对应不同的故障域和架构要求,如99%需处理节点级故障,99.9%需跨数据中心部署,99.99%需多区域冗余。

入选理由:99%可靠性需处理GPU硬件故障,依赖自动化健康检查和快速替换

FeaturedArticle#AI推理#可靠性#架构设计#故障域#Together AI英文
Together AI Blog 图标

The production platform for open-weight AI inference

Together AI Blog1714 字 (约 7 分钟)
85

Together AI推出的新推理平台使用户能高效部署和管理开放权重模型,兼顾控制、成本与性能。

入选理由:开放权重模型成本仅为封闭模型的几分之一,但质量相当。

FeaturedArticle#AI推理平台#开放权重模型#模型部署#Together AI英文
Strong results

Strong results

Aravind Srinivas(@AravSrinivas)83 字 (约 1 分钟)
85

Kimi K3在软件工程任务中性能与Claude Fable 5相当,但价格仅为后者35%。

入选理由:Kimi K3价格仅为Claude Fable 5的35%但性能相同

FeaturedTweet#模型对比#软件工程#AI性能#成本效益英文
Together AI Blog 图标

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Together AI Blog1068 字 (约 5 分钟)
85

Inkling是Thinking Machines Lab推出的多模态模型,支持高效推理和跨任务能力,Together AI提供生产级部署服务。

入选理由:Inkling通过query-conditioned attention和MoE架构实现多模态高效推理

FeaturedArticle#Inkling#多模态模型#推理平台#Together AI英文
Together AI Blog 图标

Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

Together AI Blog427 字 (约 2 分钟)
85

Together AI 获得 ISO 27001:2022 认证,证明其信息安全管理符合国际标准,增强客户对其平台安全性的信任。

入选理由:Together AI 获得 ISO 27001:2022 认证,由 A-LIGN Compliance 和 Security 颁发。

FeaturedArticle#AI#信息安全#ISO 27001#Together AI英文
Engineering voice agents: Latency, quality, and scale — Rishabh Bhargava, Together AI

Building high-quality, low-latency, scalable voice agents is now an engineering challenge requiring real-time response (<500ms), complex instruction handling, and tool calling — supported by Together AI’s infrastructure.

入选理由:语音代理必须在500毫秒内响应,否则用户会挂断电话,实时性是核心指标。

FeaturedVideo#Voice AI#Latency Optimization#Together AI#Agent Engineering英文
How Together AI built the world’s fastest speech-to-text stack

How Together AI built the world’s fastest speech-to-text stack

Together AI Blog1720 字 (约 7 分钟)
85

Together AI optimized their speech-to-text stack, achieving faster transcription speeds by using profile-aware TensorRT, optimizing the decoder loop, and improving CPU paths. They serve the two lowest-latency models, with the fastest model transcribing 20 hours of speech in under 10 seconds.

入选理由:Together AI built the world's fastest speech-to-text stack.

FeaturedArticle#Together AI#speech-to-text英文
Benchmarking inference at scale: coding agents

Benchmarking inference at scale: coding agents

Together AI Blog1358 字 (约 6 分钟)
85

Together Inference Engine delivers 31% more TPS than next fastest OSS engine on same hardware, maintains 2× better TTFT at saturation. Performance gains come from full-stack optimization.

入选理由:ThunderMLA、自定义内核重写和端到端优化使Together引擎比其他OSS引擎多31%的TPS

FeaturedArticle#Together AI#Inference Engine#Coding Agent#Performance Optimization#TTFT英文
Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

Together AI Blog979 字 (约 4 分钟)
85

Together AI and Pearl Research Labs have partnered to reduce AI inference costs through technologies like FlashAttention-4 and ATLAS.

入选理由:FlashAttention-4 提升推理速度达 1.3 倍。

FeaturedArticle#AI#Inference Optimization英文
Violin: An open-source video translation skill that breaks language barriers

Violin: An open-source video translation skill that breaks language barriers

Together AI Blog1617 字 (约 7 分钟)
75

Violin is an open-source video translation tool developed by Together AI, using multimodal models to achieve high-quality video content localization.

入选理由:Violin 支持多语言视频翻译,提升跨语言内容可访问性。

FeaturedArticle#AI#Video Processing#Natural Language Processing英文
DeepSeek-V4 Pro now available on Together AI

DeepSeek-V4 Pro Now Available on Together AI

Together AI Blog1895 字 (约 8 分钟)
75

Together AI launches DeepSeek-V4 Pro model with high-performance inference and multiple computing options.

入选理由:DeepSeek-V4 Pro 在 NVIDIA Blackwell 上实现 1.3 倍速度提升。

FeaturedArticle#AI#Model Deployment#Deep Learning中文
Foundational research powering efficient inference at scale

Foundational research powering efficient inference at scale

Together AI Blog2272 字 (约 10 分钟)
75

文章介绍了Together AI的多项技术进展,包括FlashAttention-4、ATLAS加速器和Batch Inference API更新,显著提升了大规模推理效率。

入选理由:FlashAttention-4比cuDNN快1.3倍

FeaturedArticle#AI#Inference#Efficiency#Together AI英文

跨材料问答 · Together AI

回答基于:Together AI 相关 23 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.