> ## Documentation Index
> Fetch the complete documentation index at: https://docs.archetypeai.app/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Start with /introduction/getting-started. Use the Direct Query API (POST /query) with the Newton Fusion model (text, image, and video reasoning) or the Newton Omega encoder (time-series embeddings). ATAI_API_ENDPOINT must include the version path: /v0.5 for most APIs, /v0.6 for the Fine-Tuning Service. Pages whose descriptions are marked (Archived) document the legacy Lens runtime — do not use them for new projects.

# NVIDIA VSS with Newton Fusion

> Use Newton Fusion as the video model (VLM) of NVIDIA's Video Search and Summarization (VSS) blueprint, through Newton Fusion's OpenAI-compatible endpoint.

NVIDIA's [Video Search and Summarization (VSS)](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization) blueprint answers questions about videos with an agent: an LLM plans, and a vision language model (VLM) looks at the video. VSS can call any **remote VLM** that speaks the OpenAI Chat Completions API. This guide sets up **Newton Fusion** as that VLM, through **Newton Fusion's OpenAI-compatible endpoint**.

With this setup, Newton Fusion receives the **full video with its audio**. It sees the frames and hears what is said (as a transcript), so it can answer questions about both. VSS needs configuration changes only, no code changes.

<Note>
  This example runs on an **NVIDIA DGX Spark** (GB10 Grace Blackwell, 128 GB unified memory). Newton Fusion, the VSS agent's LLM (NVIDIA Nemotron Nano 9B v2) and the rest of VSS all run on the one box, with no cloud calls at run time.
</Note>

<Warning>
  **Newton Fusion's OpenAI-compatible endpoint is not publicly distributed yet.** The Newton Fusion model weights and the software that serves them are provided by Archetype AI. To run this integration, contact [support@archetypeai.dev](mailto:support@archetypeai.dev).
</Warning>

<img src="https://mintcdn.com/archetypeai-bd6fb3cf/oDcnCaepJ0UTp8E-/images/integrations/nvidia-vss-newton-fusion.png?fit=max&auto=format&n=oDcnCaepJ0UTp8E-&q=85&s=a77d10103c6979d250bf04272a2f9d5c" alt="The VSS web UI on a DGX Spark: an assembly video on the left, and the VSS Agent's description of it on the right, produced by Newton Fusion" width="3008" height="1618" data-path="images/integrations/nvidia-vss-newton-fusion.png" />

## How it works

<img src="https://mintcdn.com/archetypeai-bd6fb3cf/oDcnCaepJ0UTp8E-/images/integrations/nvidia-vss-architecture.png?fit=max&auto=format&n=oDcnCaepJ0UTp8E-&q=85&s=b0153144ef2a5feda5c027c5aba147a7" alt="How VSS calls Newton Fusion on a DGX Spark: the VSS agent plans with Nemotron, fetches the video from VST, and sends the full MP4 with audio to Newton Fusion's OpenAI-compatible endpoint, where Whisper transcribes the audio for Newton Fusion" width="3136" height="620" data-path="images/integrations/nvidia-vss-architecture.png" />

1. You ask a question about an uploaded video.
2. The VSS agent's planner (Nemotron) decides to call its `video_understanding` tool.
3. VSS sends the video to Newton Fusion's endpoint as a Chat Completions request, with the MP4 inline.
4. Newton Fusion answers from the frames and the speech transcript. Nemotron then writes the final reply.

### Models

| Role in VSS | Model | Runs as | Memory on the DGX Spark |
| - | - | - | - |
| VLM (video understanding) | **Newton Fusion F1.0** | Newton Fusion's OpenAI-compatible endpoint | \~47 GB at 64 frames per video, Whisper included |
| Speech to text, feeding Newton Fusion | OpenAI Whisper large-v3 | Part of the same endpoint | (included above) |
| LLM (planner) | NVIDIA Nemotron Nano 9B v2 FP8 | VSS's vLLM container | \~32 GB at `--gpu-memory-utilization 0.18` |

Newton Fusion replaces VSS's default local VLM (NVIDIA Cosmos), which isn't deployed.

### Why the model is called `newton-fusion-omni-f1.0`

VSS 3.2.1 decides what to send its VLM **from the model's name**:

| Model name | What VSS sends |
| - | - |
| contains `omni`, with `ENABLE_AUDIO=true` | the full MP4, **with audio** |
| contains `cosmos` | the full MP4, no audio |
| anything else | JPEG frames only; **audio is dropped** |

Newton Fusion's endpoint therefore serves the model under the name `newton-fusion-omni-f1.0`, so VSS sends it the whole video with its sound. The weights are the same as Newton Fusion F1.0.

## Requirements

* An NVIDIA DGX Spark, with Docker and the NVIDIA container runtime (included in DGX OS)
* NVIDIA VSS **3.2.1**
* An NGC personal key with the **NGC Catalog** service, to pull the VSS containers (step 2)
* Newton Fusion's endpoint package from Archetype AI (step 0)

Steps are labelled **Workstation** (your laptop or desktop), **Spark** (a shell on the DGX Spark) or **Browser**.

## Set it up

<Steps>
  <Step stepNumber={0} title="Get Newton Fusion's endpoint (Workstation, then Spark)">
    Contact [support@archetypeai.dev](mailto:support@archetypeai.dev) for the endpoint package. It contains the serving software for the DGX Spark (arm64, CUDA), its configuration, and the Newton Fusion F1.0 and Whisper large-v3 weights (\~38 GB).

    Copy the package to the DGX Spark and follow its README to place the weights. Newton Fusion's weights and the Whisper weights are both required: Whisper is what lets Newton Fusion use the audio.
  </Step>

  <Step stepNumber={1} title="Start Newton Fusion's endpoint (Spark)">
    Start the endpoint as described in the package README. Loading the models takes about 3 minutes. The endpoint then listens on port `9095` and accepts two model names: `newton-fusion-omni-f1.0` (use this one with VSS) and `f1.0`.

    Check that it answers:

    ```bash theme={"system"}
    curl -s localhost:9095/v1/chat/completions -H 'Content-Type: application/json' \
      -d '{"model":"newton-fusion-omni-f1.0","messages":[{"role":"user","content":"Say OK."}],"max_tokens":200}'
    ```

    The first request after start-up takes up to a minute (one-time GPU setup); later ones answer in seconds.

    <Accordion title="Ask about a video directly, the way VSS does">
      VSS sends the MP4 inline, base64-encoded, as a `video_url` content part:

      ```bash theme={"system"}
      python3 - <<'EOF'
      import base64, json, urllib.request
      video = base64.b64encode(open("my_video.mp4", "rb").read()).decode()
      body = {
          "model": "newton-fusion-omni-f1.0",
          "max_tokens": 4096,
          "temperature": 0,
          "messages": [{"role": "user", "content": [
              {"type": "text", "text": "Describe what happens in this video, step by step, with times."},
              {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video}"}},
          ]}],
      }
      req = urllib.request.Request("http://localhost:9095/v1/chat/completions",
                                   json.dumps(body).encode(), {"Content-Type": "application/json"})
      print(json.load(urllib.request.urlopen(req, timeout=900))["choices"][0]["message"]["content"])
      EOF
      ```

      Requests are limited to 64 MiB, so keep videos under about 45 MB (base64 adds a third). Newton Fusion sees 64 frames per video, at low resolution, so 480p video loses nothing.
    </Accordion>
  </Step>

  <Step stepNumber={2} title="Get an NVIDIA key (Browser, then Spark)">
    VSS's containers come from NVIDIA's container registry, `nvcr.io`.

    1. Sign in at [ngc.nvidia.com](https://ngc.nvidia.com). Open your profile (top right) → **Setup** → **Generate Personal Key**.
    2. Under **Services Included**, select **NGC Catalog**. That's the only service needed here: Newton Fusion and Nemotron both run on the DGX Spark, so **Public API Endpoints** (NVIDIA-hosted models) is optional.
    3. Generate the key and copy it (`nvapi-…`). It's shown only once.

    On the DGX Spark, export it and test the login:

    ```bash theme={"system"}
    export NGC_CLI_API_KEY=nvapi-...
    export NVIDIA_API_KEY=$NGC_CLI_API_KEY     # VSS expects both; the same value is fine here
    echo "$NGC_CLI_API_KEY" | docker login nvcr.io -u '$oauthtoken' --password-stdin   # literal $oauthtoken
    ```

    `Login Succeeded` means the key works. A bare `unauthorized` means **NGC Catalog** isn't on the key.
  </Step>

  <Step stepNumber={3} title="Check for another VSS deployment (Spark)">
    VSS uses fixed container names and host ports (3000, 7777, 8000, 30000, 30001, 30081, 30888), so only one VSS deployment can run on a machine. Check for one:

    ```bash theme={"system"}
    docker ps -a --filter label=com.docker.compose.project=mdx --format '{{.Names}}\t{{.Status}}'
    ```

    <Warning>
      VSS's `dev-profile.sh up` (next step) removes an existing VSS deployment **with its volumes**, then deletes **every unused Docker volume on the host**, not only VSS's. On a machine shared with other projects, check with their owners first, and run it with `--dry-run` to see what it will do.
    </Warning>
  </Step>

  <Step stepNumber={4} title="Deploy VSS with Newton Fusion as its VLM (Spark)">
    Get VSS 3.2.1:

    ```bash theme={"system"}
    git clone --branch v3.2.1 --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
    cd video-search-and-summarization/deploy/docker
    ```

    Make two file edits that no command-line flag covers:

    ```bash theme={"system"}
    # Keep the audio: with the "omni" model name, VSS then sends the full MP4 with its sound.
    sed -i 's|^ENABLE_AUDIO=.*|ENABLE_AUDIO=true|' developer-profiles/dev-profile-base/.env

    # Give Nemotron 18% of the memory instead of 85%, leaving room for Newton Fusion.
    sed -i '/- --gpu-memory-utilization/{n;s|"0\.85"|"0.18"|}' services/nim/nvidia-nemotron-nano-9b-v2-fp8/compose.yml
    ```

    Then deploy, pointing VSS at Newton Fusion's endpoint:

    ```bash theme={"system"}
    export VLM_ENDPOINT_URL=http://172.17.0.1:9095   # the host, as seen from VSS's containers; no /v1
    export OPENAI_API_KEY=unused                     # the endpoint doesn't check it

    # --use-remote-vlm takes no value: it skips VSS's own VLM and calls VLM_ENDPOINT_URL instead
    scripts/dev-profile.sh up -p base -H DGX-SPARK -e <spark-ip> \
        --llm nvidia/NVIDIA-Nemotron-Nano-9B-v2-FP8 \
        --use-remote-vlm \
        --vlm-model-type nim \
        --vlm newton-fusion-omni-f1.0
    ```

    `<spark-ip>` is the address your browser uses to reach the DGX Spark. `172.17.0.1` is Docker's default bridge gateway; check yours with `ip -4 addr show docker0`. On first start, Nemotron downloads its weights (\~10 GB), which takes several minutes.

    Compared with NVIDIA's [remote VLM example](https://docs.nvidia.com/vss/latest/vss-agent/configure-vlm.html) (GPT-5.2 on OpenAI):

    | NVIDIA's example | For Newton Fusion | Why |
    | - | - | - |
    | `VLM_ENDPOINT_URL=https://api.openai.com` | `http://172.17.0.1:9095` | Newton Fusion's endpoint on the same DGX Spark |
    | `--vlm-model-type openai` | `--vlm-model-type nim` | In VSS 3.2.1, `openai` ignores `VLM_ENDPOINT_URL`, and `vllm` rejects the extra field VSS adds when sending audio. `nim` uses the URL and passes the request through |
    | `--vlm gpt-5.2` | `--vlm newton-fusion-omni-f1.0` | `omni` in the name, with `ENABLE_AUDIO=true`, means the full MP4 with audio |
    | `--llm-device-id 0` | `-H DGX-SPARK --llm …-FP8` | The DGX Spark profile; the FP8 Nemotron fits beside Newton Fusion |
  </Step>

  <Step stepNumber={5} title="Ask about a video (Browser, Spark)">
    Open `http://<spark-ip>:3000`. Upload a video (**+ Upload Video**), select it, click **+ Chat**, and ask, for example, *"Describe what's going on in the video."* Answers take 1–3 minutes.

    To confirm that Newton Fusion answered, on the DGX Spark:

    ```bash theme={"system"}
    docker logs vss-agent 2>&1 | grep "Using VLM profile"
    # nim_vlm, vlm_mode: remote, enable_audio: True, ..., use_video_file_base64: True

    docker logs vss-agent 2>&1 | grep "audio preserved"
    # one line per question: Downloading MP4 for inline base64 payload (audio preserved)
    ```

    `use_video_file_base64: True` means VSS sends the full MP4, not frames.
  </Step>

  <Step stepNumber={6} title="Shut it down (Spark)">
    ```bash theme={"system"}
    cd video-search-and-summarization/deploy/docker
    scripts/dev-profile.sh down --dry-run   # optional: print what it will do
    scripts/dev-profile.sh down
    ```

    Then stop Newton Fusion's endpoint as described in the package README.

    <Warning>
      `dev-profile.sh down` deletes the deployment's volumes (uploaded videos, the VST database), **every unused Docker volume on the host**, and VSS's data folder. To keep the uploads, run `docker compose -p mdx down` (without `-v`) instead.
    </Warning>
  </Step>
</Steps>

## Tips

* **Name the video and the tool in the question, and ask for one call.** Nemotron sometimes replies with a clarifying question, or calls Newton Fusion several times and rewrites its answer. This wording keeps it to one call, with Newton Fusion's answer passed through:

  > Analyze the video "\<name>" with video\_understanding. Call it once, on the whole video, with my request below word for word as its user\_prompt, and present its result as your answer without analyzing again: \<your question>

  In our tests it cut a typical answer from about 220 s to about 100 s.
* **Ask for an explicit format for time stamps**, such as *"Put each step on its own line, in the form "MM:SS - MM:SS: what happens", using the actual times from the video."* With a looser request, Newton Fusion may answer in one paragraph, and Nemotron may then invent step times when it rewrites the answer.
* **Ask about what is said, not other sounds.** Newton Fusion hears a video through its speech transcript, so narration and dialogue come through, but alarms, engines or music mostly don't.
* **Prefer Newton Fusion's answer when the two differ.** Newton Fusion answers deterministically (temperature 0). Nemotron's rewrite is where details get dropped or changed.

## Known issues

| Symptom | Cause and fix |
| - | - |
| VLM error 401 naming `platform.openai.com` | `--vlm-model-type openai` ignores `VLM_ENDPOINT_URL`. Use `nim` |
| `unexpected keyword argument 'mm_processor_kwargs'` | `--vlm-model-type vllm` with the `omni` name. Use `nim` |
| 404 on `/v1/v1/chat/completions` | `VLM_ENDPOINT_URL` ends in `/v1`. Remove it |
| 404 `model_not_found` | `--vlm` isn't `newton-fusion-omni-f1.0` or `f1.0` |
| VSS sends frames, not the MP4 | `ENABLE_AUDIO` was exported instead of set in `developer-profiles/dev-profile-base/.env` |
| Nemotron restarts with "Free memory on device … less than desired GPU memory utilization" | Its memory share is too high beside Newton Fusion. Set it to `0.18` (step 4) |
| A question about **part** of a video fails | VSS's video storage can cut partial clips whose video header can't be read up front. Archetype's endpoint package re-encodes such clips before they reach the model; if you still see this, contact support |
| The first answer after a restart is slow | Normal: one-time GPU setup. Send one request before a demo |
| An answer appears even with Newton Fusion stopped | It comes from the chat history. Test in a new chat |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.