> ## Documentation Index
> Fetch the complete documentation index at: https://docs.trymaitai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy finetune run

> Deploy a completed finetune run as a hosted model.
Promotes the finetune into a servable model addressed by `model_name`: - a new `model_name` publishes the finetune as a new model; - an existing model's `model_name` with `promote_existing: true` swaps that model's underlying finetune to this run while keeping the same model reference (so callers targeting that model name pick up the new weights).
Deployment is asynchronous: this returns immediately with the run in a `SCHEDULED` deployment state, poll `GET /finetune-runs/{run_id}/status` until `phase` reaches `deployed` (or `deploy_failed`/`deploy_cancelled`). Once deployed, run a test set against it by creating a test run whose `config.model` is this `model_name`. Hosted-model deployment is billed, so an account over its plan limit or with billing blocked returns 402.



## OpenAPI

````yaml /openapi.yaml post /finetune-runs/{run_id}/deploy
openapi: 3.1.0
info:
  title: Maitai Platform API
  description: >
    The Maitai Platform API lets you programmatically manage every resource
    available in the Maitai Portal, applications, intents, agents, datasets,
    test sets, finetune runs, and more.


    ## Authentication


    All endpoints require a valid Maitai API key passed via the
    `X-Maitai-Api-Key` header.

    You can create API keys in the [Portal](https://portal.trymaitai.com) under
    **Settings > API Keys**.
  version: 1.0.0
  contact:
    name: Maitai Support
    email: support@trymaitai.com
    url: https://trymaitai.ai
servers:
  - url: https://api.trymaitai.ai/api/v1
    description: Production
security:
  - ApiKeyAuth: []
tags:
  - name: Applications
    description: >-
      Manage applications and their configuration, intents, sessions, workflow
      runs, and models.
  - name: Analytics
    description: Read request volume and fault-rate analytics.
  - name: Intents
    description: Manage intents (application actions) nested under applications.
  - name: Intent Groups
    description: >-
      Cross-application intent grouping and access to related resources like
      sentinels, models, test sets, and datasets.
  - name: Agents
    description: >-
      Manage agents, actions, sub-agents, versions, releases, sessions, routing
      rules, and form fields.
  - name: Compositions
    description: Manage dataset compositions used for finetuning.
  - name: Sentinels
    description: >-
      Manage sentinels, evaluation watchers that monitor model output quality at
      the intent group level.
  - name: Monitors
    description: >-
      Manage reusable production monitors and their target attachments.

      Surface area summary (everything is scoped to the caller's company_id):

      * Core CRUD on monitors and their target attachments (intent / workflow /
      agent). * Lifecycle helpers (activate / pause monitor; enable / disable
      target). * Versioning + named releases (publish a snapshot, manage release
      pointers). * Activity time series + metrics rollups (charts, "is X
      healthy?" calls). * Run history (paginated `monitor_run` browsing +
      reverse lookup by source). * Live preview (run a monitor against an ad-hoc
      payload without persisting). * Discovery (find monitors attached to a
      given target). * Sample fetch (peek at the most recent production payload
      for a target).
  - name: Sessions
    description: >-
      View and manage classic (non-agent) chat sessions, timelines, and
      feedback.
  - name: Requests
    description: View and correct individual request/response pairs from chat completions.
  - name: Datasets
    description: Manage curated request sets used for model training (finetune datasets).
  - name: Evaluation Criteria
    description: Manage reusable evaluation criteria for test runs.
  - name: Unified Test Sets
    description: >-
      Unified Test Sets.

      Consumer-agnostic test-set family, the same set can drive MODEL, WORKFLOW,
      and AGENT runs. Sits alongside the legacy `test_sets.py` blueprint
      (mounted at `/test-sets`) for a graceful deprecation window; portal/SDK
      callers migrate here without breaking existing integrations.

      Sets are curated collections of `(canonical_input,
      canonical_expected_output)` pairs. Items can be hand-authored (MANUAL) or
      imported from a production `chat_completion_request`, `workflow_run`, or
      agent task; imports are idempotent at the DB level via partial unique
      indexes. When scoring a run, an optional per-item `expected_output`
      override (`test_set_item_response_sub`) shadows the item's baseline,
      that's how the portal supports "correct the answer, re-score" without
      mutating the imported ground truth.
  - name: Unified Test Runs
    description: >-
      Unified Test Runs.

      Execute a unified `test_set` against a MODEL / WORKFLOW / AGENT consumer
      and score the result. Sits alongside the legacy `test_runs.py` blueprint
      (mounted at `/test-runs`) for a graceful deprecation window; the legacy
      blueprint keeps serving the per-family run endpoints, portal/SDK callers
      migrate here.

      Lifecycle (state machine on `unified_test_run.status`):

      CREATED -> RUNNING -> COMPLETED

      * ``POST /``            -> CREATED (validates consumer_config, derives
      consumer_id, no items yet) * ``POST /{id}/prepare`` -> RUNNING (walks the
      set, adapts inputs, inserts per-item run rows) * ``POST /{id}/execute`` ->
      **non-blocking**. Schedules the conductor (in-proc asyncio task in dev,
      K8s Job in prod) and returns the run in RUNNING. Actual per-item work
      happens off-thread; callers poll ``/{id}/progress`` for completion. *
      ``GET /{id}/progress`` -> aggregate counts + duration percentiles. Portal
      polls this while a run is in flight. * ``POST /{id}/score``   -> unchanged
      status; populates matched / diff_result per item and returns a summary.
      Idempotent; safe to rerun after adding a response-sub override on an item.
      * ``POST /{id}/reset-failed`` -> COMPLETED -> RUNNING for reruns of just
      the FAILED items (S6). Skipped if the run has zero failures. * ``DELETE
      /{id}``      -> hard-delete run + items (S7).

      Plus a pure read path (``POST /preview-adapted-input``) that runs the
      input adapter for a single item against hypothetical run parameters, the
      portal wizard uses this to render "here's your adapted input" without
      paying for a real run.
  - name: Workflows
    description: >-
      Manage workflows, workflow versions/releases, artifacts, runs, and
      execution.
  - name: Finetune Runs
    description: Create, monitor, and cancel model finetuning jobs.
  - name: Evaluations
    description: >-
      Run batch evaluations against sentinels and view per-request scoring
      results.
  - name: Models
    description: View and manage available base and finetuned models.
  - name: Reports
    description: Read fault and fallback reports.
  - name: Tags
    description: Manage request tagging rules and tag runs.
  - name: Search
    description: Natural-language search over the Maitai API + CLI surface.
  - name: Help
    description: Natural-language implementation help for the Maitai platform.
paths:
  /finetune-runs/{run_id}/deploy:
    post:
      tags:
        - Finetune Runs
      summary: Deploy finetune run
      description: >-
        Deploy a completed finetune run as a hosted model.

        Promotes the finetune into a servable model addressed by `model_name`: -
        a new `model_name` publishes the finetune as a new model; - an existing
        model's `model_name` with `promote_existing: true` swaps that model's
        underlying finetune to this run while keeping the same model reference
        (so callers targeting that model name pick up the new weights).

        Deployment is asynchronous: this returns immediately with the run in a
        `SCHEDULED` deployment state, poll `GET /finetune-runs/{run_id}/status`
        until `phase` reaches `deployed` (or
        `deploy_failed`/`deploy_cancelled`). Once deployed, run a test set
        against it by creating a test run whose `config.model` is this
        `model_name`. Hosted-model deployment is billed, so an account over its
        plan limit or with billing blocked returns 402.
      operationId: deployFinetuneRun
      parameters:
        - name: run_id
          in: path
          required: true
          schema:
            type: integer
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              properties:
                model_name:
                  type: string
                  description: >-
                    Name for the deployed hosted model. Use this same value as
                    `config.model` when running a test set against it.
                promote_existing:
                  type: boolean
                  x-cli-option: true
                  description: >-
                    Reuse an already-provisioned instance if one is available
                    instead of provisioning fresh.
              required:
                - model_name
            example:
              model_name: support-sft-v1
              promote_existing: false
      responses:
        '200':
          description: Deploy finetune run.
          content:
            application/json:
              schema:
                type: object
                properties:
                  data:
                    $ref: '#/components/schemas/FinetuneRun'
              example:
                data:
                  id: 61
                  status: COMPLETED
                  meta:
                    model_deployment_status: SCHEDULED
components:
  schemas:
    FinetuneRun:
      type: object
      description: A model finetuning job that trains a new model on a dataset.
      properties:
        id:
          type: integer
          readOnly: true
          description: Unique identifier.
        company_id:
          type: integer
          readOnly: true
          description: Owning company ID.
        application_id:
          type: integer
          nullable: true
          description: Application scope.
        intent_id:
          type: integer
          nullable: true
          description: Intent scope.
        intent_group_id:
          type: integer
          nullable: true
          description: Intent group scope.
        date_created:
          type: string
          format: date-time
          readOnly: true
          description: When the job was created (UTC).
        base_model:
          type: string
          description: Base model being finetuned (e.g. `gpt-4o-mini-2024-07-18`).
        display_name:
          type: string
          nullable: true
          description: Human-readable name for the run.
        status:
          type: string
          nullable: true
          description: >-
            Job status: `PENDING`: queued, `TRAINING`: in progress, `COMPLETED`:
            done, `FAILED`: errored, `CANCELLED`: stopped.
          enum:
            - PENDING
            - TRAINING
            - COMPLETED
            - FAILED
            - CANCELLED
        progress:
          type: number
          format: double
          description: Training progress from 0.0 to 1.0.
        meta:
          type: object
          nullable: true
          description: Hyperparameters, output model info, and other job metadata.
  securitySchemes:
    ApiKeyAuth:
      type: apiKey
      in: header
      name: X-Maitai-Api-Key
      description: Your Maitai API key from the Portal.

````