> ## Documentation Index
> Fetch the complete documentation index at: https://docs.archetypeai.app/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Start with /introduction/getting-started. Use the Direct Query API (POST /query) with the Newton Fusion model (text, image, and video reasoning) or the Newton Omega encoder (time-series embeddings). ATAI_API_ENDPOINT must include the version path: /v0.5 for most APIs, /v0.6 for the Fine-Tuning Service. Pages whose descriptions are marked (Archived) document the legacy Lens runtime — do not use them for new projects.

# Get Eval

> Retrieve one eval, its status, and its metrics report

<Callout icon="clock" color="#3064E3" iconType="solid">
  Requires [version 1.1.12](/release-notes/1.1.x#v1-1-12) or later of the Archetype platform.
</Callout>

## Overview

This endpoint returns one eval by its `evl_` id.

Fields beyond the always-present ones are populated as the eval progresses:

* `started_at` is populated when the runner picks up the eval
* `output_artifacts` is populated as the run progresses. An entry whose
  `metadata.status` is `partial` holds the rows scored so far, which is also what a failed or
  cancelled mid-run eval leaves behind.
* `metrics_report` is populated upon successful completion and is `null` until then
* `completed_at` is populated upon successful completion
* `error` is populated on a `failed` completion.

The `metrics_report` answers the question "is this agent good enough?" To answer "where is this
agent bad?" page through [List Eval Examples](/api-reference/agents/evals/list-eval-examples).

## Request

<ParamField path="eval_id" type="string" required>
  Eval `evl_` id.
</ParamField>

## Response

<ResponseField name="id" type="string" required>
  TypeID-encoded eval identifier (`evl_` prefix).
</ResponseField>

<ResponseField name="name" type="string" required>
  Human label for the eval.
</ResponseField>

<ResponseField name="org_id" type="string" required>
  Organization identifier the eval belongs to.
</ResponseField>

<ResponseField name="blueprint_id" type="string" required>
  The blueprint being evaluated.
</ResponseField>

<ResponseField name="primary" type="string" required>
  The run's headline, as `<target>.<objective>` — resolved at creation, so this is what was
  actually scored rather than what was asked for.
</ResponseField>

<ResponseField name="examples" type="array" required>
  The examples the eval was created against, as resolved: each carries the name its results are
  keyed by and inputs stamped with the CRC32C of the bytes scored and the fully-resolved
  ground-truth declarations.

  <Note>
    `examples` records what the run actually read; that is, resolved names, the CRC32C of the
    bytes scored, and fully-resolved ground-truth declarations, rather than what the request asked
    for.
  </Note>
</ResponseField>

<ResponseField name="rendered_config" type="object" required>
  The fully-expanded configuration this eval runs under, resolved when the eval was created. An
  eval does not run the blueprint quite unchanged: it runs the whole pipeline with a scoring
  stage in place of the blueprint's own sink, and the ground-truth labels declared alongside each
  input. Opaque JSON on the wire — treat the shape as informational.
</ResponseField>

<ResponseField name="status" type="string" required>
  Eval lifecycle status: `pending`, `running`, `completed`, `failed`, or `cancelled`.

  <Note>
    Evals are always created with the status `pending`. Once they're dispatched, their state
    changes to `running`. From there, it will eventually enter one of the three terminal states:
    `completed`, `failed`, or `cancelled`.
  </Note>
</ResponseField>

<ResponseField name="created_by" type="string" required>
  Subject id (`usr_...` or `key_...`) that created this eval.
</ResponseField>

<ResponseField name="created_at" type="string" required>
  Creation timestamp (date-time).
</ResponseField>

<ResponseField name="started_at" type="string">
  When the runner picked the eval up; `null` before then.
</ResponseField>

<ResponseField name="completed_at" type="string">
  When the eval finished; `null` while unfinished.
</ResponseField>

<ResponseField name="metrics_report" type="object">
  Aggregate and per-target scoring results. Populated when the eval reaches `completed`; `null`
  while pending, running, cancelled, or failed. See [Metrics
  report](#metrics-report-metrics_report) for the format of this object.
</ResponseField>

<ResponseField name="output_artifacts" type="array">
  Refs to the files the run produced. Each `id` is a data-service file id, so an artifact is
  downloadable through the files API, and each carries `metadata` saying which artifact it is.
  See [Output artifact ref](#output-artifact-ref-output_artifacts) for the format of the objects
  in this array.

  Populated *while the run goes*, not only at the end: each batch of scored rows is appended to
  the predictions file and updates the entry, whose `metadata.status` stays `partial` until the
  run's final push marks it `complete`. An eval that failed or was cancelled mid-run keeps its
  partial entry — the rows scored so far are still readable.
</ResponseField>

<ResponseField name="error" type="string">
  Failure detail; `null` unless the eval failed.
</ResponseField>

### Metrics report (`metrics_report`)

The run's headline plus one entry per scored target.

<ResponseField name="schema_version" type="string" required>
  Wire-format version. Bumped on breaking changes; additive changes (a new target type, a new
  optional field) stay on the same version.
</ResponseField>

<ResponseField name="primary" type="object" required>
  The run's headline score. Deliberately repeats a value that also appears in the named target's
  `aggregate`: a reader wanting the one number the run is judged by gets it without following a
  pointer into the map.

  * `target` (string, required) — the target this headline is about.
  * `name` (string, required) — the metric's name, e.g. `macro_f1`.
  * `value` (number, required) — its value.
</ResponseField>

<ResponseField name="targets" type="object" required>
  One entry per scored target, keyed by target name. See [Target
  report](#target-report-metrics_report-targets-\<target-name>) for the format of a target
  record.
</ResponseField>

### Target report (`metrics_report.targets.<target name>`)

<ResponseField name="aggregate" type="object" required>
  Every objective this run computed for the target, by name.
</ResponseField>

<ResponseField name="type" type="string" required>
  The target's type. For `category` — one class per scored unit — the three fields below are also present.
</ResponseField>

<ResponseField name="class_names" type="array" required>
  The class vocabulary, fixing the index order of `per_class` and the confusion matrix.
</ResponseField>

<ResponseField name="per_class" type="object" required>
  Per-class `precision`, `recall`, `f1`, and `support` — parallel arrays aligned to
  `class_names`. `support` counts real occurrences of each class in the ground-truth stream,
  independent of what the classifier predicted.
</ResponseField>

<ResponseField name="confusion_matrix" type="array" required>
  Row = true class, column = predicted class, both in `class_names` order.
</ResponseField>

### Output artifact ref (`output_artifacts[]`)

<ResponseField name="type" type="string" required>
  Storage kind, e.g. `file`.
</ResponseField>

<ResponseField name="id" type="string" required>
  Data-service file id, so the artifact is downloadable through the files API.
</ResponseField>

<ResponseField name="format" type="string">
  Optional format hint.
</ResponseField>

<ResponseField name="crc32c" type="string">
  Whole-file CRC32C checksum of the referenced bytes, base64-encoded exactly as S3 emits it.
  `null` when the data service has no checksum for the file.
</ResponseField>

<ResponseField name="metadata" type="object">
  For an artifact an eval produced: `kind` (which artifact it is), `status` (`partial` or
  `complete`), and `row_count` (rows in the file — how many scored rows a reader will find in it).
</ResponseField>

<RequestExample>
  ```bash cURL theme={"system"}
  curl "$ATAI_API_URL/agents/evals/evl_01jcb0h2m6t4xr9nv3k7pdzs5y" \
    -H "Authorization: Bearer $ATAI_API_KEY"
  ```

  ```python Python theme={"system"}
  import os
  import time

  import requests

  base_url = os.environ["ATAI_API_URL"]
  api_key = os.environ["ATAI_API_KEY"]
  headers = {"Authorization": f"Bearer {api_key}"}

  eval_id = "evl_01jcb0h2m6t4xr9nv3k7pdzs5y"

  while True:
      response = requests.get(f"{base_url}/agents/evals/{eval_id}", headers=headers)
      evaluation = response.json()

      if evaluation["status"] in ("completed", "failed", "cancelled"):
          break
      time.sleep(5)

  if evaluation["status"] == "completed":
      headline = evaluation["metrics_report"]["primary"]
      print(f"{headline['target']}.{headline['name']} = {headline['value']}")
  else:
      print(f"{evaluation['status']}: {evaluation.get('error')}")
  ```

  ```javascript JavaScript theme={"system"}
  const response = await fetch(
    `${process.env.ATAI_API_URL}/agents/evals/evl_01jcb0h2m6t4xr9nv3k7pdzs5y`,
    {
      headers: {
        'Authorization': `Bearer ${process.env.ATAI_API_KEY}`
      }
    }
  );

  const body = await response.json();

  if (response.ok) {
    console.log(`${body.name} [${body.status}]`);
    if (body.metrics_report) {
      const { target, name, value } = body.metrics_report.primary;
      console.log(`${target}.${name} = ${value}`);
    }
  } else {
    console.error('Error:', body.errors);
  }
  ```
</RequestExample>

<ResponseExample>
  ```json 200 - Completed theme={"system"}
  {
    "id": "evl_01jcb0h2m6t4xr9nv3k7pdzs5y",
    "name": "Pump A regression",
    "org_id": "org_01jc8m5r2vq9xt4bn7h3kdzs6w",
    "blueprint_id": "blp_01jc9n7k3xf8mbq2v5t0ary6de",
    "primary": "state.macro_f1",
    "examples": [
      {
        "name": "site-a-morning",
        "ordinal": 1,
        "inputs": [
          {"type": "file", "id": "file_abc123", "format": "csv", "crc32c": "AAAAAA=="}
        ],
        "metadata": {"site": "a", "shift": "morning"}
      }
    ],
    "rendered_config": {},
    "status": "completed",
    "metrics_report": {
      "schema_version": "v1",
      "primary": {"target": "state", "name": "macro_f1", "value": 0.87},
      "targets": {
        "state": {
          "type": "category",
          "aggregate": {"macro_f1": 0.87, "accuracy": 0.91},
          "class_names": ["running", "idle", "fault_bearing"],
          "per_class": {
            "precision": [0.93, 0.88, 0.74],
            "recall": [0.96, 0.85, 0.69],
            "f1": [0.94, 0.86, 0.71],
            "support": [4120, 1880, 260]
          },
          "confusion_matrix": [
            [3955, 140, 25],
            [220, 1598, 62],
            [45, 36, 179]
          ]
        }
      }
    },
    "output_artifacts": [
      {
        "type": "file",
        "id": "file_jkl012",
        "format": "ndjson",
        "crc32c": "AAAAAA==",
        "metadata": {"kind": "predictions", "status": "complete", "row_count": 6260}
      }
    ],
    "created_by": "usr_01jc8m4p3rt6vx9qn2h5kdzb7y",
    "created_at": "2026-09-18T11:24:03Z",
    "started_at": "2026-09-18T11:24:09Z",
    "completed_at": "2026-09-18T11:31:52Z",
    "error": null
  }
  ```

  ```json 404 - Eval not found theme={"system"}
  {
    "errors": [
      {
        "code": "<error_code>",
        "message": "Eval not found.",
        "suggestion": null,
        "error_uid": "err-xxxxxxxx"
      }
    ]
  }
  ```
</ResponseExample>

## Important Notes

<Note>
  The run-level figure is every example pooled, never the mean of the per-example results from
  [List Eval Examples](/api-reference/agents/evals/list-eval-examples).
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.