Skip to main content
GET
Get run status

Authorizations

Authorization
string
header
required

MCPJam API key (sk_…). Create one at Settings → API keys. Guest sessions cannot use the API, and API keys cannot manage other API keys.

Path Parameters

projectId
string
required

ID of the hosted project that contains the server.

runId
string
required

Eval run ID, as returned by POST /eval-runs.

Response

The run.

id
string
required
suiteId
string
required
status
enum<string>
required

Poll until terminal: completed, failed, or cancelled.

Available options:
pending,
running,
completed,
failed,
cancelled
source
enum<string>
required

Run origin. API-created runs are api.

Available options:
ui,
api,
sdk
createdAt
number
required

Epoch milliseconds.

runNumber
integer | null
result
enum<string> | null

Verdict once terminal. inconclusive exists only under verdictPolicyVersion: 2 and is NOT a failure: the run did not measure the server well enough to say (too few gradeable trials, too many evaluator errors), so a gate that folds it into failed reports a defect the run never observed. Read verdictSummary.reasons for the check that withheld the verdict.

Available options:
passed,
failed,
inconclusive,
null
summary
object | null
notes
string | null
completedAt
number | null

Epoch milliseconds, null until terminal.

scoreIntegrity
enum<string> | null

Whether the run's score evidence verified at ingest. TRI-STATE, and the third state matters: valid means the backend checked and definitions and results agree; invalid means they do not; null (or absent) means NO VERDICT was produced, on a deployment that predates integrity checking. A score gate must treat null exactly like invalid — absent evidence is not valid evidence.

Available options:
valid,
invalid,
null
verdictPolicyVersion
enum<integer>

The verdict policy this run was decided under, frozen at run start. ABSENT means legacy percent-threshold grading — result cannot then be inconclusive and there is no verdictSummary. A caller gating on fractions or on validity must read this FIRST rather than assume a missing summary means a clean run.

Available options:
2
verdictSummary
object

How a policy-2 verdict was reached: the resolved validity policy, the measured completion and evaluator-error rates with their denominators and exclusions, the per-case and per-execution-variant aggregates, and the exact reasons. Absent when the run is legacy, or when the stored summary failed contract validation at the boundary — a partially-valid decision is never published, because a gate cannot tell a missing field from a satisfied check.

verdictPolicyIntegrityError
string

Why a policy-2 run could not be decided from its own evidence (a missing or malformed policy snapshot, mixed evaluator configs). Accompanies an inconclusive result; it is never a task failure.

environment
object · null · null

The environment revision this run is pinned to. null on a legacy run that recorded none — always present, so a caller never has to distinguish absent from unpinned.

runGroupId
string

Shared by every per-target run from the same fan-out launch. Absent on a single-target launch and on rows created before run groups.

effectiveModelId
string

Model the run actually executed with. Absent on pre-attribution rows.

modelSource
enum<string>

client_default inherited the host model; override used the environment's modelId.

Available options:
client_default,
override
executionEngine
string

Which engine executed the run: emulated (the platform's own turn loop) or harness:<id> (a real agent runtime such as Claude Code). ABSENT means the run recorded no engine — a run created before the platform attributed one. Treat that as UNKNOWN, never as emulated: those are different claims, and the runs whose engine was never recorded are exactly the ones a reader must not vouch for.

insights
object

The common actionable-insights envelope. Present on the DETAIL response only — lists stay compact — and absent when the caller may not have it or the deployment cannot produce one. Treat absence exactly like status: "not_available".

judges
object

Advisory LLM graders on this run. Present on the DETAIL response only — lists stay compact — and absent on deployments that predate the envelope.

gateWaiver
object | null

The waiver currently in force over this run's gate, or null. Gated on being able to VIEW the run, deliberately not on being able to grant a waiver — a waiver only its grantors can see is not a visible one. Carried on the run so a client computing its own gate can fold a waiver in and name it, without a second round trip.

importEligibility
object

Whether a run's imported cases carry evidence a gate may rely on. Computed by the platform from the run's OWN frozen snapshot, never from the suite's current cases — recomputing from those would let a later edit change what a finished run is allowed to prove. legacy means the run contains no imported cases at all (every native run, forever) and is gateable unchanged; eligible means every imported case carries a valid frozen decision; incomplete means the evidence cannot be trusted, which makes the run NOT GATEABLE and is not a test verdict. The whole field is ABSENT on deployments that predate import eligibility — a different fact from legacy, and one a gate must read as "no opinion, behave as before".