Together AI Blog

Deploy and inference any model from HuggingFace

7.5内容质量
Deploy and inference any model from HuggingFace

TL;DR · AI 摘要

文章介绍了如何使用Together AI平台部署和推理任何来自HuggingFace的模型,提供了多种部署选项和硬件支持。

核心要点

  • Together AI平台支持部署和推理HuggingFace模型
  • 提供服务器无感推理、批量推理等多种选项
  • 支持NVIDIA多种GPU型号

结构提纲

按章节快速跳转。

  1. 介绍Together AI平台及其支持的功能

  2. 列出服务器无感推理、批量推理等部署选项

  3. 支持NVIDIA多种GPU型号,如GB300、GB200等

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • Together AI平台部署和推理HuggingFace模型

金句 / Highlights

值得收藏与分享的关键句。

#Together AI#HuggingFace#模型部署#推理
打开原文

Deploy and inference any model from HuggingFace

Image 3⚡️ FlashAttention-4: up to 1.3× faster than cuDNN on NVIDIA Blackwell →

Image 4Introducing Together AI's new look →

Image 5🔎 ATLAS: runtime-learning accelerators delivering up to 4x faster LLM inference →

Image 6⚡ Together GPU Clusters: self-service NVIDIA GPUs, now generally available →

Image 7📦 Batch Inference API: Process billions of tokens at 50% lower cost for most models →

Image 8🪛 Fine-Tuning Platform Upgrades: Larger Models, Longer Contexts →

[](https://www.together.ai/)

  • ![Image 9 Serverless Inference High-performance inference as APIs](https://www.together.ai/serverless-inference)
  • ![Image 10 Batch Inference Inference for batch workloads](https://www.together.ai/batch-inference)
  • ![Image 11 Dedicated Model Inference Inference on custom hardware](https://www.together.ai/dedicated-model-inference)
  • ![Image 12 Dedicated Container Inference Inference for custom models](https://www.together.ai/dedicated-container-inference)

![Image 13 MiniMax M2.5 Image 14 Nano Banana Pro Image 15 Qwen3.5-397B Image 16 GLM-5 Image 17 kimi k2.5 Image 18 gpt-oss-120B Model library Explore the top open-source models](https://www.together.ai/models)

Accelerated Compute

  • ![Image 19 GPU Clusters Reliable GPU clusters at scale](https://www.together.ai/gpu-clusters)
  • ![Image 20 AI Factory Custom infrastructure at frontier scale](https://www.together.ai/ai-factory)

Developer Environments

  • ![Image 21 Sandbox Build development environments for AI](https://www.together.ai/sandbox)

Storage

  • ![Image 22 Managed Storage Store model weights & data securely](https://www.together.ai/managed-storage)
  • ![Image 23 Fine-Tuning Shape models with your data](https://www.together.ai/fine-tuning)
  • ![Image 24 Evaluations Measure model quality](https://www.together.ai/evaluations)

![Image 25 DeepSeek V3.1 Image 26 GLM 5 FP4 Image 27 Qwen3-VL 32B Image 28 gpt-oss-120b Image 29 kimi k2.5 Image 30 Llama 4 Maverick Model library Fine-tune top open-source models](https://www.together.ai/models)

  • ![Image 31 Research Systems research for production AI](https://www.together.ai/research)
  • ![Image 32 Research blog All our research publications](https://www.together.ai/research-blog)

Featured publications

Show all

  • ![Image 33 Documentation Technical docs for Together AI](https://docs.together.ai/)
  • ![Image 34 Demos Our open-source demo apps](https://www.together.ai/demos)
  • ![Image 35 Cookbooks Practical implementation guides](https://www.together.ai/cookbooks)
  • ![Image 36 Voice Agents Build voice agents for production](https://www.together.ai/solutions/voice)

Resources

  • ![Image 37 Customer stories Testimonials from AI Natives](https://www.together.ai/customers)
  • ![Image 38 Startup accelerator Build and scale your startup](https://www.together.ai/startup-accelerator)
  • ![Image 39 Customer support Find answers to your questions](https://www.together.ai/support)
  • ![Image 40 Blog Our latest news & blog posts](https://www.together.ai/blog)
  • ![Image 41 Events Explore our events calendar](https://www.together.ai/events)

Company

  • ![Image 42 About Get to know us](https://www.together.ai/about-us)
  • ![Image 43 Careers Join our mission](https://www.together.ai/careers)

*

  • ![Image 44 Serverless Inference High-performance inference as APIs](https://www.together.ai/serverless-inference)
  • ![Image 45 Batch Inference Inference for batch workloads](https://www.together.ai/batch-inference)
  • ![Image 46 Dedicated Model Inference Inference on custom hardware](https://www.together.ai/dedicated-model-inference)
  • ![Image 47 Dedicated Container Inference Inference for custom models](https://www.together.ai/dedicated-container-inference)

![Image 48 MiniMax M2.5 Image 49 Nano Banana Pro Image 50 Qwen3.5-397B Image 51 GLM-5 Image 52 kimi k2.5 Image 53 gpt-oss-120B Model library Explore the top open-source models](https://www.together.ai/models)

* Accelerated Compute

  • ![Image 54 GPU Clusters Reliable GPU clusters at scale](https://www.together.ai/gpu-clusters)
  • ![Image 55 AI Factory Custom infrastructure at frontier scale](https://www.together.ai/ai-factory)

Developer Environments

  • ![Image 56 Sandbox Build development environments for AI](https://www.together.ai/sandbox)

Storage

  • ![Image 57 Managed Storage Store model weights & data securely](https://www.together.ai/managed-storage)

*

  • ![Image 58 Fine-Tuning Shape models with your data](https://www.together.ai/fine-tuning)
  • ![Image 59 Evaluations Measure model quality](https://www.together.ai/evaluations)

![Image 60 DeepSeek V3.1 Image 61 GLM 5 FP4 Image 62 Qwen3-VL 32B Image 63 gpt-oss-120b Image 64 kimi k2.5 Image 65 Llama 4 Maverick Model library Fine-tune top open-source models](https://www.together.ai/models)

*

  • ![Image 66 Research Systems research for production AI](https://www.together.ai/research)
  • ![Image 67 Research blog All our research publications](https://www.together.ai/research-blog)

Featured publications

Show all

*

  • ![Image 68 Documentation Technical docs for Together AI](https://docs.together.ai/)
  • ![Image 69 Demos Our open-source demo apps](https://www.together.ai/demos)
  • ![Image 70 Cookbooks Practical implementation guides](https://www.together.ai/cookbooks)
  • ![Image 71 Voice Agents Build voice agents for production](https://www.together.ai/solutions/voice)

* Resources

  • ![Image 72 Customer stories Testimonials from AI Natives](https://www.together.ai/customers)
  • ![Image 73 Startup accelerator Build and scale your startup](https://www.together.ai/startup-accelerator)
  • ![Image 74 Customer support Find answers to your questions](https://www.together.ai/support)
  • ![Image 75 Blog Our latest news & blog posts](https://www.together.ai/blog)
  • ![Image 76 Events Explore our events calendar](https://www.together.ai/events)

Company

  • ![Image 77 About Get to know us](https://www.together.ai/about-us)
  • ![Image 78 Careers Join our mission](https://www.together.ai/careers)

Contact sales

Contact sales

Sign in

All blog posts

Inference

Published 5/8/2026

Deploy and inference any model from HuggingFace

Agents, Skills, and Together Dedicated Container Inference unlocks trying any model.

  • Authors Blaine Kasten
  • Table of contents

Something real is shifting in how developers work. Agents open up work that used to be off-limits, not because it was technically impossible, but because it required niche expertise most of us didn't have. Containerization, inference server configs, model-specific environment setup: these are the kinds of tasks that used to demand either deep expertise or hours of self-education before you could even get started. Agents allow for an elegant way to bridge those pre-requisite knowledge gaps. You describe what you want, and the agent fills in the knowledge gaps.

That's the unlock. Not speed. _Access._‍

The day Netflix dropped a new model

Netflix recently released void-model on Hugging Face. The day it came out, my instinct was the same as always: I want to try this. But wanting to try a new model and actually _running_ it are two different things. Getting it into a usable environment, handling the inference server setup, figuring out the container configuration, wiring it all up correctly: that's the part that usually introduces a day or two of lag between "this looks cool" and "okay I'm actually using it."

This time, that lag was basically zero.

Using Goose, a CLI agent runner, combined with Together's dedicated containers skill, I went from "Netflix just dropped a model" to "I have a running container for it" in a single session. The agent produced all the code needed to deploy void-model on Together's Dedicated Container Inference (DCI) infrastructure, essentially on release day.

The output lives here: github.com/blainekasten/together-void-model-container

Exactly what I did

The whole setup took three steps.

Step 1: Install the Together dedicated containers skill.

npx skills add togethercomputer/skills

Image 79
Image 79

That pulls in the together-dedicated-containers skill, which gives Goose the specific knowledge it needs to work with Together's infrastructure: how to configure the inference server, what the container spec should look like, how to wire everything up for a given model.

Step 2: Start a Goose session and run one prompt.

I want to deploy this model on togethers dedicated containers https://huggingface.co/netflix/void-model

Image 80
Image 80

That's it. One sentence.

Step 3: Sit back and watch it work.

From there, the agent pulled the model details from Hugging Face, figured out the right inference server configuration for the model architecture, generated the container config files, and produced a complete, runnable setup, all without me having to look anything up or guide it through individual steps.

The result: blainekasten/together-void-model-container, a clean, working repo anyone can use to run void-model on Together infrastructure.

Step 4: Use your model!

After the agent deploys your application you can start running inference against it. The Together CLIhas commands to easily test inference.

bash
tg beta jig submit --watch --payload '{
    "video_url": "https://github.com/Netflix/void-model/raw/refs/heads/main/sample/lime/input_video.mp4",
    "quadmask_url": "https://github.com/Netflix/void-model/raw/refs/heads/main/sample/lime/quadmask_0.mp4",
    "prompt": "Empty park bench with fallen leaves on the ground",
    "use_pass2": false
  }'

This model removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed.

Our inference calls with this model are asynchronous. Therefore the response of this request will return a payload with an identifier we can poll for. The response looks like this:

bash
{
  "model": "void-byoc",
  "request_id": "019dc0f3-3c73-7a3f-b4b6-87ad06091180",
  "status": "running",
  "claimed_at": "2026-04-24T19:24:19.447457Z",
  "created_at": "2026-04-24T19:24:19.444567Z",
  "done_at": null,
  "info": null,
  "inputs": {
    "prompt": "Empty park bench with fallen leaves on the ground",
    "quadmask_url": "https://github.com/Netflix/void-model/raw/refs/heads/main/sample/lime/quadmask_0.mp4",
    "use_pass2": false,
    "video_url": "https://github.com/Netflix/void-model/raw/refs/heads/main/sample/lime/input_video.mp4"
  },
  "outputs": null,
  "priority": 1,
  "retries": null,
  "warnings": null
}

When the inference completes, the outputs includes a URL to the hosted video. We can download it using cURL and our Together API key:

bash
curl -L -O \
  https://api.together.ai/v1/storage/019dc0f3-3c73-7a3f-b4b6-87ad06091180-tmpddmhtvar.mp4 \
  --header "Authorization: Bearer $TOGETHER_API_KEY"

Note: -L is required to follow the http redirect in the storage url and -O will write the output to a local file.

Why Together Dedicated Container Inference

This story only works because Together's Dedicated Container Inference (DCI) is genuinely a great place to run models like this, and it's worth explaining why.

DCI gives you a private, GPU-backed environment running the model of your choice, fully managed by Together. You're not fighting for shared resources, you're not configuring your own cluster, and you're not locked into a fixed menu of available models. You bring the model; Together handles the infrastructure.

This is a big deal for teams that want to move fast. When a new model drops from Netflix, from a research lab, from the open-source community, you can have it running in a production-grade environment almost immediately. No spinning up your own GPU VMs, no wrestling with inference server dependencies, no waiting for someone to add support for it in a managed endpoint. DCI is flexible by design: if the model exists, you can deploy it.

The cost model also makes it easy to experiment. You're paying for what you use, on a container that's yours, without the overhead of managing the underlying compute. That's the kind of setup that lets you say yes to testing new models instead of filing it away for "when I have time."

If you're interested in Together's DCI, reach out to us to get set up.

Start building on Together AI

From optimized training and model shaping to large-scale production inference

Get Started now

Image 81
Image 81

* Products

  • Models

See all modelsDeepSeek Meta Qwen Google OpenAI Mistral AI Custom models * Developers

Pricing

* Resources

© 2026 Together AI. All Rights Reserved.

  • [](https://discord.gg/9Rk6sSeWEG)
  • [](https://x.com/togethercompute)
  • [](https://www.linkedin.com/company/togethercomputer/)

Image 83Image 84