> ## Documentation Index
> Fetch the complete documentation index at: https://mcpjam-claude-project-secrets-implementation-e3rb5i.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate eval cases from the suite's tools

> Discovers the suite's server tools over a live MCP connection, generates cases against them, and persists them — the only edit route that connects to a server, and the only one that SPENDS ORG CREDITS. Synchronous: connect, generate, persist, disconnect, respond.

An environment-based suite generates against that environment's closed server set, so the cases match the tools its runs will actually see.

Pass `x-mcpjam-idempotency-key` to make a retry safe: drafts are recorded backend-side BEFORE any case is persisted, so a replay reuses them instead of spending credits again, and each case is persisted under a derived per-item key so the loop is resumable.



## OpenAPI

````yaml /reference/openapi.json post /projects/{projectId}/eval-suites/{suiteId}/cases/generate
openapi: 3.1.0
info:
  title: MCPJam API
  version: 1.0.0-preview
  description: >-
    Programmatic access to MCP servers saved in your MCPJam projects — live
    diagnostics (validate, inspect, export) and operations: call tools, render
    prompts, run eval suites asynchronously and poll their results, and import
    OAuth tokens.


    **The API is in preview**: the surface may change while we finish the
    design. Error `code` values are stable; error `message` strings are not.
    Write clients that ignore unknown response fields.
  contact:
    name: MCPJam
    url: https://github.com/MCPJam/inspector/issues
servers:
  - url: https://app.mcpjam.com/api/v1
    description: Hosted MCPJam
security:
  - bearerAuth: []
tags:
  - name: Clients
    description: >-
      Clients — the named, reusable configurations that define how MCPJam
      connects to and talks to your MCP servers. The original `/hosts` paths
      remain as deprecated, ID-only compatibility aliases with their original
      DTOs and their original (tokenless) write contracts; every alias response
      carries `Deprecation: true`. New integrations should use `/clients`.
  - name: Environments
    description: >-
      Project environments: named, live-editable execution bundles (one host, an
      optional standalone server group, optionally pinned skills and plugin
      versions) that eval suites and journeys run against. Distinct from Sandbox
      images, which are Computer base images. Reads require project membership;
      every write requires project admin.
  - name: Plugins
    description: >-
      Agent Plugins imported into a project — read-only inventory and version
      detail.
  - name: Skills
    description: >-
      Cloud Skills: authored SKILL.md files stored in a project. Read-only here.
      Environments pin skills by id (`skillSelection.skillIds`) and eval runs
      pin them with `--compose-skill`, so this surface exists to give an
      unattended caller those ids; authoring is an app flow behind a beta gate.
  - name: Sandbox images
    description: >-
      Custom Computer images: a digest-pinned Dockerfile built into an immutable
      image your project's computers boot from.
  - name: Server diagnostics
    description: Connect-level health checks against a saved MCP server.
  - name: Primitives
    description: 'The server''s MCP primitives: tools, prompts, and resources.'
  - name: Export
    description: Full-server snapshots for diffing and CI.
  - name: Execution
    description: 'Run the server''s primitives: call tools, render prompts.'
  - name: Eval runs
    description: >-
      Asynchronous eval suite runs: create with 202, poll status, iterations,
      and traces.
  - name: Conformance runs
    description: >-
      Ingest MCP spec-conformance results from the SDK/CLI into project-owned
      history. Distinct from Eval runs (authored LLM cases) and from directory
      readiness.
  - name: Server connections
    description: >-
      Connect an MCP server URL to a project, authorizing in a browser when the
      server requires it.
  - name: OAuth
    description: 'Bring-your-own OAuth: import externally obtained tokens for a server.'
  - name: Scenarios
    description: >-
      Read-only access to the scenarios published from a project: listing,
      settings, attached servers, and share links.
  - name: Catalog
    description: >-
      Discover the resources the other routes operate on: your account,
      projects, servers, eval suites, and chat sessions.
  - name: Tunnels
    description: >-
      Relay tunnels that expose local MCP servers through a public URL,
      registered as first-class project servers (the `mcpjam cloud tunnel` CLI
      flow).
  - name: Agent
    description: >-
      Headless agent turns over the public API: send a message history, the
      server runs one assistant turn with project-scoped workspace tools (eval
      reads + suite creation) on a pinned hosted model, and returns the reply
      plus created-resource references.
  - name: Swarms
    description: >-
      Personas, journeys and swarm containers — the authoring half of Swarms —
      plus the model-backed generation that drafts them.
  - name: Swarm runs
    description: >-
      Launching journeys and reading what they produced. Launching SPENDS — see
      the per-operation notes.
  - name: Swarm insights
    description: >-
      What a swarm run revealed. The scorecard and findings are deterministic
      and free; requesting wave insights runs models and draws on your shared
      daily ledger.
  - name: User testing
    description: >-
      Publishing an environment for real visitors, and controlling who can reach
      it. Several of these NARROW access and take effect immediately.
  - name: Directory readiness
    description: >-
      Grade a saved server against a publisher's listing requirements:
      Anthropic's connector directory or OpenAI's plugin directory. Reported as
      lane status and coverage, never as a numeric score, and excluded from
      `pooledConformanceScore`. Deterministic grading is free; model-backed
      experience observations are an explicit opt-in that consumes MCPJam
      credits and can never decide a verdict.
  - name: Registry
    description: >-
      Search the scraped MCP directories (Claude, ChatGPT, and any future
      source), list curated/org registry cards, and install them into a project.
      Install writes a `servers` row and provenance — it does not open a live
      session. There is no catalog-uninstall route: delete the project server
      instead. Directory reads require a bearer (including minted guest tokens)
      but do not materialize a user. Card/connection reads and all writes are
      authed-non-guest.
paths:
  /projects/{projectId}/eval-suites/{suiteId}/cases/generate:
    post:
      tags:
        - Eval runs
      summary: Generate eval cases from the suite's tools
      description: >-
        Discovers the suite's server tools over a live MCP connection, generates
        cases against them, and persists them — the only edit route that
        connects to a server, and the only one that SPENDS ORG CREDITS.
        Synchronous: connect, generate, persist, disconnect, respond.


        An environment-based suite generates against that environment's closed
        server set, so the cases match the tools its runs will actually see.


        Pass `x-mcpjam-idempotency-key` to make a retry safe: drafts are
        recorded backend-side BEFORE any case is persisted, so a replay reuses
        them instead of spending credits again, and each case is persisted under
        a derived per-item key so the loop is resumable.
      operationId: generateEvalCases
      parameters:
        - $ref: '#/components/parameters/projectId'
        - $ref: '#/components/parameters/suiteId'
        - name: x-mcpjam-idempotency-key
          in: header
          required: false
          schema:
            type: string
          description: >-
            Makes a retry replay the recorded drafts instead of spending credits
            again.
      requestBody:
        required: false
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvalCaseGenerateRequest'
      responses:
        '200':
          description: The generated cases.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvalCaseGenerated'
        '400':
          $ref: '#/components/responses/ValidationError'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
        '409':
          $ref: '#/components/responses/Conflict'
        '429':
          $ref: '#/components/responses/RateLimited'
        '500':
          $ref: '#/components/responses/InternalError'
        '502':
          $ref: '#/components/responses/ServerUnreachable'
        '504':
          $ref: '#/components/responses/Timeout'
components:
  parameters:
    projectId:
      name: projectId
      in: path
      required: true
      description: ID of the hosted project that contains the server.
      schema:
        type: string
    suiteId:
      name: suiteId
      in: path
      required: true
      description: Eval suite ID, as returned by `POST /eval-runs`.
      schema:
        type: string
  schemas:
    EvalCaseGenerateRequest:
      type: object
      description: >-
        AI-generate cases from the suite's server tools and persist them. SPENDS
        ORG CREDITS.
      properties:
        mode:
          type: string
          enum:
            - normal
            - negative
          description: Superseded by `caseMix` when that is present.
        servers:
          type: array
          items:
            type: string
            minLength: 1
          description: >-
            Server ids or names to discover tools from. Ignored when the suite
            is environment-based.
        environmentId:
          type: string
          minLength: 1
          description: >-
            Discover tools from this attached environment's closed server set,
            so generated cases are written against the tools the suite's runs
            will actually see.
        caseModels:
          type: array
          items:
            type: object
            required:
              - model
            properties:
              model:
                type: string
              provider:
                type: string
        caseMix:
          type: object
          description: >-
            Per-bucket case counts. Omitted buckets inherit the default mix; the
            backend bounds each bucket and the total.
          properties:
            simple:
              type: integer
              minimum: 0
              maximum: 10
            multiTool:
              type: integer
              minimum: 0
              maximum: 10
            multiTurn:
              type: integer
              minimum: 0
              maximum: 10
            complex:
              type: integer
              minimum: 0
              maximum: 10
            negative:
              type: integer
              minimum: 0
              maximum: 10
        varyUserStyles:
          type: boolean
          description: >-
            Condition generated cases on a range of user styles so the queries
            read like different users wrote them.
        idempotencyKey:
          type: string
          minLength: 1
          maxLength: 256
          description: >-
            Write-idempotency key. A repeat call with the same key replays
            recorded drafts instead of spending credits again. The
            `Idempotency-Key` / `x-mcpjam-idempotency-key` header carries the
            same value and WINS over this field.
      additionalProperties: false
    EvalCaseGenerated:
      type: object
      required:
        - generationModel
        - created
        - counts
      properties:
        generationModel:
          type: string
        created:
          type: array
          items:
            $ref: '#/components/schemas/EvalCase'
        counts:
          type: object
          properties:
            normal:
              type: integer
            negative:
              type: integer
        skipped:
          type: array
          description: >-
            Drafts that were generated but failed to persist. Surfaced rather
            than silently dropped.
          items:
            type: object
    EvalCase:
      type: object
      required:
        - id
        - title
        - steps
        - iterations
        - isNegative
        - models
      description: >-
        A persisted eval case, in the public steps-first shape. Note this is NOT
        `EvalTestCase`, which is the INLINE authoring shape accepted by suite
        creation.
      properties:
        id:
          type: string
        declaredId:
          type: string
          description: >-
            The case's effective declared id. Absent on cases authored before
            declared identity existed.
        title:
          type: string
        steps:
          type: array
          minItems: 1
          items:
            $ref: '#/components/schemas/EvalTestStep'
          description: >-
            Ordered test steps. A `prompt` step is a model turn; a single
            model-free `toolCall` step is a render-check; `assert` steps hold
            the expectations.
        expectedOutput:
          type: string
        iterations:
          type: integer
          minimum: 1
          maximum: 10
        repetitions:
          type: integer
          minimum: 1
          description: >-
            Trials this case runs under verdict policy 2, overriding the suite
            default. Absent means the case inherits it. NOT a second spelling of
            `iterations`: that one is the legacy count, which the legacy
            resolver reads as a FLOOR (`max(iterations,
            suite.minimumIterations)`) and which a policy-2 case still reports
            for compatibility. This one is exact.
        passThreshold:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Fraction of this case's trials that must pass, overriding the suite
            default. Absent means the case inherits it. Never derived from the
            suite's `minimumAccuracy`, which is a PERCENT under a different
            resolver.
        isNegative:
          type: boolean
          description: When true, the case passes if NO tools are called.
        scenario:
          type: string
        intent:
          type: string
          minLength: 1
          maxLength: 64
          pattern: ^\S(?:[\s\S]*\S)?$
          description: >-
            Optional authored analytics grouping label. Must be already trimmed;
            absent means unlabelled.
        models:
          type: array
          items:
            type: object
            required:
              - model
            properties:
              model:
                type: string
              provider:
                type: string
        matchOptions:
          type: object
          description: >-
            Absent when the case sets none — omitted from the response rather
            than sent as `null`.
        checks:
          type: object
          description: >-
            Absent when the case sets none — omitted from the response rather
            than sent as `null`.
          properties:
            mode:
              type: string
              enum:
                - inherit
                - replace
                - extend
            list:
              type: array
              items:
                type: object
        import:
          $ref: '#/components/schemas/EvalCaseImportClaim'
        createdAt:
          type:
            - number
            - 'null'
        updatedAt:
          type:
            - number
            - 'null'
    Error:
      type: object
      required:
        - code
        - message
      properties:
        code:
          type: string
          description: >-
            Stable, machine-readable error code. New codes may be added over
            time; treat unknown codes as non-retryable failures unless the HTTP
            status says otherwise.
          enum:
            - UNAUTHORIZED
            - FORBIDDEN
            - NOT_FOUND
            - CONFLICT
            - VALIDATION_ERROR
            - RATE_LIMITED
            - FEATURE_NOT_SUPPORTED
            - SERVER_UNREACHABLE
            - TIMEOUT
            - OAUTH_REQUIRED
            - INTERNAL_ERROR
        message:
          type: string
          description: >-
            Human-readable description. May change between releases — don't
            match on it.
        details:
          type: object
          description: Optional, unstructured context bag.
          additionalProperties: true
    EvalTestStep:
      type: object
      description: >-
        One authored test step (the unified test model). `kind` discriminates:
        `prompt` is a user message (model turn); `toolCall` is a deterministic,
        model-free tool call; `interact` is one pure widget action; `assert` is
        an assertion (a `Predicate` like `toolCalledWith` / `widgetRendered`, or
        a DOM `WidgetAssertion`).
      required:
        - id
        - kind
      properties:
        id:
          type: string
          minLength: 1
        kind:
          type: string
          enum:
            - prompt
            - toolCall
            - interact
            - assert
        prompt:
          type: string
          description: 'User message (`kind: prompt`).'
        serverName:
          type: string
          description: 'Server that owns the tool (`kind: toolCall`).'
        toolName:
          type: string
          description: 'Tool name (`kind: toolCall` / `interact`).'
        arguments:
          type: object
          description: 'Tool-call arguments (`kind: toolCall`).'
          additionalProperties: true
        action:
          type: object
          description: 'Widget action (`kind: interact`).'
          additionalProperties: true
        assertion:
          type: object
          description: 'Predicate or widget assertion (`kind: assert`).'
          additionalProperties: true
      additionalProperties: true
    EvalCaseImportClaim:
      type: object
      description: >-
        What a converter CLAIMED about one imported case. `exact` is
        CONVERTER-CLAIMED exact — the converter says it applied a structural
        mapping rule, cited in `note`. MCPJam has NOT verified semantic
        equivalence, so user-facing copy must say "claimed exact", never
        "verified" or "accepted". Claim-only: who approved an approximation,
        when, and why is a PER-RUN decision frozen on the run
        (`ImportEligibility.approvedApproximationReceipts`), never stored on the
        case — an approval that lived on a case would outlive the run it was
        granted for and the edit that invalidated it. Approval and internal keys
        are rejected with 400, never stripped.
      required:
        - status
      additionalProperties: false
      properties:
        status:
          type: string
          enum:
            - exact
            - approximated
            - unsupported
            - unresolved
          description: >-
            `exact`: the converter claims a 1:1 structural mapping, and must
            cite it in `note`. `approximated`: behaviour was intentionally
            approximated; a human must approve it for EVERY run. `unsupported`:
            the source behaviour cannot currently be represented. `unresolved`:
            a deterministic reference does not resolve against the live target.
            A selected `unsupported` or `unresolved` case cannot run.
        sourceCaseKey:
          type: string
          minLength: 1
          maxLength: 512
          description: The case's identity in the source system, when it had one.
        note:
          type: string
          minLength: 1
          maxLength: 2000
          description: >-
            Why the status is what it is — the mapping rule cited, or what was
            lost. REQUIRED when `status` is `exact`.
  responses:
    ValidationError:
      description: Malformed body or parameters.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: VALIDATION_ERROR
            message: Invalid JSON body
    Unauthorized:
      description: >-
        Missing, invalid, revoked, or orphaned key (`UNAUTHORIZED`) — or the
        **target MCP server** needs an OAuth grant (`OAUTH_REQUIRED`), which is
        a property of the server, not your key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          examples:
            badKey:
              summary: Invalid or revoked key
              value:
                code: UNAUTHORIZED
                message: Invalid API key
            oauthRequired:
              summary: Target server needs an OAuth grant
              value:
                code: OAUTH_REQUIRED
                message: Server requires OAuth authorization
    Forbidden:
      description: Key is valid but not allowed to do this.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: FORBIDDEN
            message: You do not have access to this project
    NotFound:
      description: Unknown project, server, or resource.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: NOT_FOUND
            message: Server not found
    Conflict:
      description: >-
        The resource is not in a state that accepts this write — a stale
        `expectedRevision`, a duplicate name, or an environment that cannot
        currently be launched. The request was well-formed; re-read the resource
        and retry.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: CONFLICT
            message: >-
              Environment changed since you loaded it (expected revision 3,
              current 5). Reload and retry.
    RateLimited:
      description: >-
        Per-key rate limit exceeded (60 requests/minute sustained, bursts up to
        10). Honor `Retry-After` and back off with jitter.
      headers:
        Retry-After:
          description: Seconds to wait before retrying.
          schema:
            type: integer
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: RATE_LIMITED
            message: API key rate limit exceeded. Slow down and retry.
    InternalError:
      description: Something failed on MCPJam's side.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: INTERNAL_ERROR
            message: Unexpected internal error
    ServerUnreachable:
      description: Could not connect to the target MCP server.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: SERVER_UNREACHABLE
            message: Failed to connect to server
    Timeout:
      description: The target MCP server connected but didn't respond in time.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: TIMEOUT
            message: Request to server timed out
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        MCPJam API key (`sk_…`). Create one at [Settings → API
        keys](https://app.mcpjam.com/settings/api-keys). Guest sessions cannot
        use the API, and API keys cannot manage other API keys.

````