Skip to main content
NVIDIA’s Video Search and Summarization (VSS) blueprint answers questions about videos with an agent: an LLM plans, and a vision language model (VLM) looks at the video. VSS can call any remote VLM that speaks the OpenAI Chat Completions API. This guide sets up Newton Fusion as that VLM, through Newton Fusion’s OpenAI-compatible endpoint. With this setup, Newton Fusion receives the full video with its audio. It sees the frames and hears what is said (as a transcript), so it can answer questions about both. VSS needs configuration changes only, no code changes.
This example runs on an NVIDIA DGX Spark (GB10 Grace Blackwell, 128 GB unified memory). Newton Fusion, the VSS agent’s LLM (NVIDIA Nemotron Nano 9B v2) and the rest of VSS all run on the one box, with no cloud calls at run time.
Newton Fusion’s OpenAI-compatible endpoint is not publicly distributed yet. The Newton Fusion model weights and the software that serves them are provided by Archetype AI. To run this integration, contact [email protected].
The VSS web UI on a DGX Spark: an assembly video on the left, and the VSS Agent's description of it on the right, produced by Newton Fusion

How it works

How VSS calls Newton Fusion on a DGX Spark: the VSS agent plans with Nemotron, fetches the video from VST, and sends the full MP4 with audio to Newton Fusion's OpenAI-compatible endpoint, where Whisper transcribes the audio for Newton Fusion
  1. You ask a question about an uploaded video.
  2. The VSS agent’s planner (Nemotron) decides to call its video_understanding tool.
  3. VSS sends the video to Newton Fusion’s endpoint as a Chat Completions request, with the MP4 inline.
  4. Newton Fusion answers from the frames and the speech transcript. Nemotron then writes the final reply.

Models

Newton Fusion replaces VSS’s default local VLM (NVIDIA Cosmos), which isn’t deployed.

Why the model is called newton-fusion-omni-f1.0

VSS 3.2.1 decides what to send its VLM from the model’s name: Newton Fusion’s endpoint therefore serves the model under the name newton-fusion-omni-f1.0, so VSS sends it the whole video with its sound. The weights are the same as Newton Fusion F1.0.

Requirements

  • An NVIDIA DGX Spark, with Docker and the NVIDIA container runtime (included in DGX OS)
  • NVIDIA VSS 3.2.1
  • An NGC personal key with the NGC Catalog service, to pull the VSS containers (step 2)
  • Newton Fusion’s endpoint package from Archetype AI (step 0)
Steps are labelled Workstation (your laptop or desktop), Spark (a shell on the DGX Spark) or Browser.

Set it up

0

Get Newton Fusion's endpoint (Workstation, then Spark)

Contact [email protected] for the endpoint package. It contains the serving software for the DGX Spark (arm64, CUDA), its configuration, and the Newton Fusion F1.0 and Whisper large-v3 weights (~38 GB).Copy the package to the DGX Spark and follow its README to place the weights. Newton Fusion’s weights and the Whisper weights are both required: Whisper is what lets Newton Fusion use the audio.
1

Start Newton Fusion's endpoint (Spark)

Start the endpoint as described in the package README. Loading the models takes about 3 minutes. The endpoint then listens on port 9095 and accepts two model names: newton-fusion-omni-f1.0 (use this one with VSS) and f1.0.Check that it answers:
The first request after start-up takes up to a minute (one-time GPU setup); later ones answer in seconds.
VSS sends the MP4 inline, base64-encoded, as a video_url content part:
Requests are limited to 64 MiB, so keep videos under about 45 MB (base64 adds a third). Newton Fusion sees 64 frames per video, at low resolution, so 480p video loses nothing.
2

Get an NVIDIA key (Browser, then Spark)

VSS’s containers come from NVIDIA’s container registry, nvcr.io.
  1. Sign in at ngc.nvidia.com. Open your profile (top right) → Setup → Generate Personal Key.
  2. Under Services Included, select NGC Catalog. That’s the only service needed here: Newton Fusion and Nemotron both run on the DGX Spark, so Public API Endpoints (NVIDIA-hosted models) is optional.
  3. Generate the key and copy it (nvapi-…). It’s shown only once.
On the DGX Spark, export it and test the login:
Login Succeeded means the key works. A bare unauthorized means NGC Catalog isn’t on the key.
3

Check for another VSS deployment (Spark)

VSS uses fixed container names and host ports (3000, 7777, 8000, 30000, 30001, 30081, 30888), so only one VSS deployment can run on a machine. Check for one:
VSS’s dev-profile.sh up (next step) removes an existing VSS deployment with its volumes, then deletes every unused Docker volume on the host, not only VSS’s. On a machine shared with other projects, check with their owners first, and run it with --dry-run to see what it will do.
4

Deploy VSS with Newton Fusion as its VLM (Spark)

Get VSS 3.2.1:
Make two file edits that no command-line flag covers:
Then deploy, pointing VSS at Newton Fusion’s endpoint:
<spark-ip> is the address your browser uses to reach the DGX Spark. 172.17.0.1 is Docker’s default bridge gateway; check yours with ip -4 addr show docker0. On first start, Nemotron downloads its weights (~10 GB), which takes several minutes.Compared with NVIDIA’s remote VLM example (GPT-5.2 on OpenAI):
5

Ask about a video (Browser, Spark)

Open http://<spark-ip>:3000. Upload a video (+ Upload Video), select it, click + Chat, and ask, for example, “Describe what’s going on in the video.” Answers take 1–3 minutes.To confirm that Newton Fusion answered, on the DGX Spark:
use_video_file_base64: True means VSS sends the full MP4, not frames.
6

Shut it down (Spark)

Then stop Newton Fusion’s endpoint as described in the package README.
dev-profile.sh down deletes the deployment’s volumes (uploaded videos, the VST database), every unused Docker volume on the host, and VSS’s data folder. To keep the uploads, run docker compose -p mdx down (without -v) instead.

Tips

  • Name the video and the tool in the question, and ask for one call. Nemotron sometimes replies with a clarifying question, or calls Newton Fusion several times and rewrites its answer. This wording keeps it to one call, with Newton Fusion’s answer passed through:
    Analyze the video “<name>” with video_understanding. Call it once, on the whole video, with my request below word for word as its user_prompt, and present its result as your answer without analyzing again: <your question>
    In our tests it cut a typical answer from about 220 s to about 100 s.
  • Ask for an explicit format for time stamps, such as “Put each step on its own line, in the form “MM:SS - MM:SS: what happens”, using the actual times from the video.” With a looser request, Newton Fusion may answer in one paragraph, and Nemotron may then invent step times when it rewrites the answer.
  • Ask about what is said, not other sounds. Newton Fusion hears a video through its speech transcript, so narration and dialogue come through, but alarms, engines or music mostly don’t.
  • Prefer Newton Fusion’s answer when the two differ. Newton Fusion answers deterministically (temperature 0). Nemotron’s rewrite is where details get dropped or changed.

Known issues