> ## Documentation Index
> Fetch the complete documentation index at: https://docs.co-mind.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create chat completion (stateless)

> OpenAI-compatible chat completions endpoint. **Stateless** — the client sends the full conversation history in `messages[]` on every request; the server does not persist prior turns. For stateful, tool-orchestrated conversations, see the Chat Sessions (Preview) endpoints (upcoming release).




## OpenAPI

````yaml /openapi.yaml post /v1/chat/completions
openapi: 3.1.0
info:
  title: Co-mind.ai Private AI Platform API
  version: 1.2.4
  description: >
    Co-mind.ai Private AI Platform API.


    ## Features

    - OpenAI-compatible endpoints

    - Tool/function calling support

    - Streaming responses

    - Multiple backend support

    - Personal Access Tokens (PAT) for programmatic access


    ## Authentication

    The API supports two authentication methods:


    **1. JWT Authentication** — Login with email/password to get short-lived
    access tokens.

    Use for interactive sessions (web apps, Postman).


    **2. Personal Access Tokens (PAT)** — Long-lived tokens for programmatic/API
    access.

    Create via `POST /v1/api-tokens` after authenticating with JWT.

    PAT format: `cmnd_<tokenId>.<secret>`


    Both methods use the `Authorization: Bearer <token>` header.
  contact:
    email: support@co-mind.ai
servers:
  - url: http://co-mind-platform-host
    description: Co-mind.ai AI Platform
security:
  - BearerAuth: []
tags:
  - name: Discovery
    description: List available models
  - name: Authentication
    description: >-
      JWT login, refresh, logout, SSO, registration, password reset, and
      `/v1/auth/me`
  - name: API Tokens
    description: Personal Access Token (PAT) management
  - name: Completions
    description: Text completion endpoints (single-shot, no conversation state)
  - name: Chat
    description: >
      OpenAI-compatible chat completions. **Stateless** — the client sends the
      full conversation history on every request; the server does not persist
      messages.
  - name: Responses
    description: >
      OpenAI-Responses-compatible endpoint. Wire-shape passthrough to the
      Inference Engine — every field the caller sends flows through, so the
      caller can use the full Responses API surface. The server does not persist
      responses server-side; use the `/v1/responses/{resp_id}/continue` variant
      to extend a prior response by id.
  - name: Embeddings
    description: >
      Text-embedding endpoint. Runs against the platform's configured embedding
      model (see `GET /v1/models`); the resulting vectors are returned to the
      caller and are not persisted server-side.
  - name: Translation
    description: >
      Text translation between languages. Ships with `POST /v1/translate/text`;
      async document translation is on the roadmap for a future release.
  - name: Knowledge Base
    description: >
      Knowledge base management (create / update / delete KBs, upload files,
      query for context) and **stateless** retrieval-augmented chat
      (`/v1/knowledgebase/chat/completions`). The RAG chat endpoint is stateless
      — the client sends the full conversation history on every request; the
      server injects retrieved KB context and returns the completion.
  - name: Chat Sessions (Preview)
    description: >
      **Preview — not currently enabled.**


      **Stateful** chat sessions. The server manages conversation history and
      supports tool orchestration (knowledge base retrieval, research, agents).
      These endpoints are documented for integrator preparation but are not
      exposed on any deployment today; they will be enabled in an upcoming
      release. Request and response schemas may change before general
      availability.
  - name: Researcher
    description: Web search, research sessions, analysis, synthesis, and source credibility
  - name: Quota
    description: Usage quota tracking
paths:
  /v1/chat/completions:
    post:
      tags:
        - Chat
      summary: Create chat completion (stateless)
      description: >
        OpenAI-compatible chat completions endpoint. **Stateless** — the client
        sends the full conversation history in `messages[]` on every request;
        the server does not persist prior turns. For stateful, tool-orchestrated
        conversations, see the Chat Sessions (Preview) endpoints (upcoming
        release).
      operationId: createChatCompletion
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatCompletionRequest'
      responses:
        '200':
          description: Successful response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ChatCompletionResponse'
        '400':
          $ref: '#/components/responses/BadRequest'
        '429':
          $ref: '#/components/responses/RateLimitExceeded'
components:
  schemas:
    ChatCompletionRequest:
      type: object
      required:
        - model
        - messages
      properties:
        model:
          type: string
          example: llama3.2:3b
        messages:
          type: array
          items:
            $ref: '#/components/schemas/ChatMessage'
        temperature:
          type: number
          minimum: 0
          maximum: 2
          default: 1
        max_tokens:
          type: integer
          minimum: 1
        top_p:
          type: number
          minimum: 0
          maximum: 1
          default: 1
        stream:
          type: boolean
          default: false
        tools:
          type: array
          items: a9baed6f-ce90-4f20-9c43-e3a8fd16bf08
    ChatCompletionResponse:
      type: object
      properties:
        id:
          type: string
          example: chatcmpl-abc123
        object:
          type: string
          example: chat.completion
        created:
          type: integer
        model:
          type: string
        choices:
          type: array
          items:
            type: object
            properties:
              index:
                type: integer
              message:
                $ref: '#/components/schemas/ChatMessage'
              finish_reason:
                type: string
                enum:
                  - stop
                  - length
                  - tool_calls
                  - content_filter
        usage:
          $ref: '#/components/schemas/Usage'
    ChatMessage:
      type: object
      required:
        - role
        - content
      properties:
        role:
          type: string
          enum:
            - system
            - user
            - assistant
            - tool
        content:
          type: string
        name:
          type: string
    Usage:
      type: object
      properties:
        prompt_tokens:
          type: integer
        completion_tokens:
          type: integer
        total_tokens:
          type: integer
    Error:
      type: object
      properties:
        error:
          type: object
          properties:
            message:
              type: string
            type:
              type: string
            code:
              type: string
  responses:
    BadRequest:
      description: Bad request
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error:
              message: Invalid request parameters
              type: invalid_request_error
              code: bad_request
    RateLimitExceeded:
      description: Rate limit exceeded
      headers:
        Retry-After:
          schema:
            type: integer
          description: Seconds to wait before retrying
        x-ratelimit-limit-requests:
          schema:
            type: integer
        x-ratelimit-remaining-requests:
          schema:
            type: integer
        x-ratelimit-reset-requests:
          schema:
            type: integer
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            error:
              message: Rate limit exceeded
              type: rate_limit_exceeded
              code: rate_limit_exceeded
  securitySchemes:
    BearerAuth:
      type: http
      scheme: bearer
      bearerFormat: JWT or PAT
      description: >
        Bearer token authentication. Supports two token types:

        - **JWT Access Token** — obtained via `POST /v1/auth/login`

        - **Personal Access Token (PAT)** — created via `POST /v1/api-tokens`,
        format: `cmnd_<tokenId>.<secret>`

````