Skip to main content
GET
Requires version 1.1.12 or later of the Archetype platform.

Overview

This endpoint returns one eval by its evl_ id. Fields beyond the always-present ones are populated as the eval progresses:
  • started_at is populated when the runner picks up the eval
  • output_artifacts is populated as the run progresses. An entry whose metadata.status is partial holds the rows scored so far, which is also what a failed or cancelled mid-run eval leaves behind.
  • metrics_report is populated upon successful completion and is null until then
  • completed_at is populated upon successful completion
  • error is populated on a failed completion.
The metrics_report answers the question “is this agent good enough?” To answer “where is this agent bad?” page through List Eval Examples.

Request

string
required
Eval evl_ id.

Response

string
required
TypeID-encoded eval identifier (evl_ prefix).
string
required
Human label for the eval.
string
required
Organization identifier the eval belongs to.
string
required
The blueprint being evaluated.
string
required
The run’s headline, as <target>.<objective> — resolved at creation, so this is what was actually scored rather than what was asked for.
array
required
The examples the eval was created against, as resolved: each carries the name its results are keyed by and inputs stamped with the CRC32C of the bytes scored and the fully-resolved ground-truth declarations.
examples records what the run actually read; that is, resolved names, the CRC32C of the bytes scored, and fully-resolved ground-truth declarations, rather than what the request asked for.
object
required
The fully-expanded configuration this eval runs under, resolved when the eval was created. An eval does not run the blueprint quite unchanged: it runs the whole pipeline with a scoring stage in place of the blueprint’s own sink, and the ground-truth labels declared alongside each input. Opaque JSON on the wire — treat the shape as informational.
string
required
Eval lifecycle status: pending, running, completed, failed, or cancelled.
Evals are always created with the status pending. Once they’re dispatched, their state changes to running. From there, it will eventually enter one of the three terminal states: completed, failed, or cancelled.
string
required
Subject id (usr_... or key_...) that created this eval.
string
required
Creation timestamp (date-time).
string
When the runner picked the eval up; null before then.
string
When the eval finished; null while unfinished.
object
Aggregate and per-target scoring results. Populated when the eval reaches completed; null while pending, running, cancelled, or failed. See Metrics report for the format of this object.
array
Refs to the files the run produced. Each id is a data-service file id, so an artifact is downloadable through the files API, and each carries metadata saying which artifact it is. See Output artifact ref for the format of the objects in this array.Populated while the run goes, not only at the end: each batch of scored rows is appended to the predictions file and updates the entry, whose metadata.status stays partial until the run’s final push marks it complete. An eval that failed or was cancelled mid-run keeps its partial entry — the rows scored so far are still readable.
string
Failure detail; null unless the eval failed.

Metrics report (metrics_report)

The run’s headline plus one entry per scored target.
string
required
Wire-format version. Bumped on breaking changes; additive changes (a new target type, a new optional field) stay on the same version.
object
required
The run’s headline score. Deliberately repeats a value that also appears in the named target’s aggregate: a reader wanting the one number the run is judged by gets it without following a pointer into the map.
  • target (string, required) — the target this headline is about.
  • name (string, required) — the metric’s name, e.g. macro_f1.
  • value (number, required) — its value.
object
required
One entry per scored target, keyed by target name. See Target report for the format of a target record.

Target report (metrics_report.targets.<target name>)

object
required
Every objective this run computed for the target, by name.
string
required
The target’s type. For category — one class per scored unit — the three fields below are also present.
array
required
The class vocabulary, fixing the index order of per_class and the confusion matrix.
object
required
Per-class precision, recall, f1, and support — parallel arrays aligned to class_names. support counts real occurrences of each class in the ground-truth stream, independent of what the classifier predicted.
array
required
Row = true class, column = predicted class, both in class_names order.

Output artifact ref (output_artifacts[])

string
required
Storage kind, e.g. file.
string
required
Data-service file id, so the artifact is downloadable through the files API.
string
Optional format hint.
string
Whole-file CRC32C checksum of the referenced bytes, base64-encoded exactly as S3 emits it. null when the data service has no checksum for the file.
object
For an artifact an eval produced: kind (which artifact it is), status (partial or complete), and row_count (rows in the file — how many scored rows a reader will find in it).

Important Notes

The run-level figure is every example pooled, never the mean of the per-example results from List Eval Examples.