Cancel a run
Request cancellation of an in-flight run; marks the run and its pending/running iterations cancelled. A no-op success when the run is already cancelled; returns 409 when the run already reached a terminal status (completed/failed/timed_out).
Authorizations
MCPJam API key (sk_…). Create one at Settings → API keys. Guest sessions cannot use the API, and API keys cannot manage other API keys.
Path Parameters
ID of the hosted project that contains the server.
Eval run ID, as returned by POST /eval-runs.
Response
The cancelled run.
Poll until terminal: completed, failed, or cancelled.
pending, running, completed, failed, cancelled Run origin. API-created runs are api.
ui, api, sdk Epoch milliseconds.
Verdict once terminal. inconclusive exists only under verdictPolicyVersion: 2 and is NOT a failure: the run did not measure the server well enough to say (too few gradeable trials, too many evaluator errors), so a gate that folds it into failed reports a defect the run never observed. Read verdictSummary.reasons for the check that withheld the verdict.
passed, failed, inconclusive, null Epoch milliseconds, null until terminal.
Whether the run's score evidence verified at ingest. TRI-STATE, and the third state matters: valid means the backend checked and definitions and results agree; invalid means they do not; null (or absent) means NO VERDICT was produced, on a deployment that predates integrity checking. A score gate must treat null exactly like invalid — absent evidence is not valid evidence.
valid, invalid, null The verdict policy this run was decided under, frozen at run start. ABSENT means legacy percent-threshold grading — result cannot then be inconclusive and there is no verdictSummary. A caller gating on fractions or on validity must read this FIRST rather than assume a missing summary means a clean run.
2 How a policy-2 verdict was reached: the resolved validity policy, the measured completion and evaluator-error rates with their denominators and exclusions, the per-case and per-execution-variant aggregates, and the exact reasons. Absent when the run is legacy, or when the stored summary failed contract validation at the boundary — a partially-valid decision is never published, because a gate cannot tell a missing field from a satisfied check.
Why a policy-2 run could not be decided from its own evidence (a missing or malformed policy snapshot, mixed evaluator configs). Accompanies an inconclusive result; it is never a task failure.
The environment revision this run is pinned to. null on a legacy run that recorded none — always present, so a caller never has to distinguish absent from unpinned.
Shared by every per-target run from the same fan-out launch. Absent on a single-target launch and on rows created before run groups.
Model the run actually executed with. Absent on pre-attribution rows.
client_default inherited the host model; override used the environment's modelId.
client_default, override Which engine executed the run: emulated (the platform's own turn loop) or harness:<id> (a real agent runtime such as Claude Code). ABSENT means the run recorded no engine — a run created before the platform attributed one. Treat that as UNKNOWN, never as emulated: those are different claims, and the runs whose engine was never recorded are exactly the ones a reader must not vouch for.
The common actionable-insights envelope. Present on the DETAIL response only — lists stay compact — and absent when the caller may not have it or the deployment cannot produce one. Treat absence exactly like status: "not_available".
Advisory LLM graders on this run. Present on the DETAIL response only — lists stay compact — and absent on deployments that predate the envelope.
The waiver currently in force over this run's gate, or null. Gated on being able to VIEW the run, deliberately not on being able to grant a waiver — a waiver only its grantors can see is not a visible one. Carried on the run so a client computing its own gate can fold a waiver in and name it, without a second round trip.
Whether a run's imported cases carry evidence a gate may rely on. Computed by the platform from the run's OWN frozen snapshot, never from the suite's current cases — recomputing from those would let a later edit change what a finished run is allowed to prove. legacy means the run contains no imported cases at all (every native run, forever) and is gateable unchanged; eligible means every imported case carries a valid frozen decision; incomplete means the evidence cannot be trusted, which makes the run NOT GATEABLE and is not a test verdict. The whole field is ABSENT on deployments that predate import eligibility — a different fact from legacy, and one a gate must read as "no opinion, behave as before".

