Skip to main content
GET
Requires version 1.1.12 or later of the Archetype platform.

Overview

This endpoint lists what an eval measured for each of its examples. The eval’s own metrics_report answers “is this agent good enough”; these answer “where is it bad”. Each example is scored on its own rows as the run streams, and the run-level report is those examples pooled — so the two always agree, and a run-level figure is never the mean of what this returns. Results are ordered by the example’s position in the run, oldest first. That is deliberately not the newest-first order the other lists use: an example set has no time order, so the useful order is the one the caller supplied. An eval that has not completed lists no examples rather than erroring. Results appear when the run reports them.
The HTTP status code returned is 404 both if the specified eval ID is unknown and if the ID exists but belongs to another organization.

Request

string
required
Eval evl_ id.
integer
default:"100"
Page size. Minimum 1, maximum 1000.
string
Forward cursor: return examples after this page’s last one. Pass the previous page’s next_cursor. Mutually exclusive with before.
string
Backward cursor: return examples before this page’s first one. Pass the current page’s prev_cursor. Mutually exclusive with after.

Response

array
required
The page, in the order the eval was given its examples. Unlike every other list in this API this is not newest-first: an example set has no time order, and reversing the list the caller supplied would only make it harder to read.Each entry in this array is an Example result object.
boolean
required
true when more results exist beyond this page in the direction of travel.
string
Cursor for the next page in the same direction — pass it as after when paging forward, or as before when you supplied before. null when has_more is false.
string
Cursor to step back the way this page was reached. null on the first page.

Example result object

string
required
Identifier of this example result.
string
required
How the example is identified — its own name, or the stem of its input file when it declared none.
integer
required
The example’s 1-based position in the run, and the order results are listed in.
array
required
The example’s input refs, each carrying the CRC32C of the bytes scored.
string
required
Whether this example was scored: completed or failed. Narrower than the eval’s own status on purpose — an example is not dispatched, paused, or cancelled on its own. It is part of a run that either got far enough to score it or did not.
object
One entry per scored target. Absent when the example was not scored. Same shape as the run report’s targets, narrowed to this example’s rows — so a client that renders the run’s numbers renders an example’s with no second code path.
targets is absent on an example whose status is failed; when the status is failed, read error instead.
object
The free-form tags the example was created with. What a long per-example table is sliced by — which site, which shift, which operator.
string
Why the example was not scored. Present when status is failed.