---
title: "RealtimeServerEventConversationItemInputAudioTranscriptionCompleted"
url: "https://kongair.terwilligar.com/apis/openai-api-2-3-0/versions/ea92d048-c746-4ead-a6a8-da28d93137d1/schemas/RealtimeServerEventConversationItemInputAudioTranscriptionCompleted"
---

> Full API specification: https://kongair.terwilligar.com/apis/openai-api-2-3-0/versions/ea92d048-c746-4ead-a6a8-da28d93137d1.md

# RealtimeServerEventConversationItemInputAudioTranscriptionCompleted

This event is the output of audio transcription for user audio written to the user audio buffer. Transcription begins when the input audio buffer is committed by the client or server (in `server_vad` mode). Transcription runs asynchronously with Response creation, so this event may come before or after the Response events. Realtime API models accept audio natively, and thus input transcription is a separate process run on a separate ASR (Automatic Speech Recognition) model, currently always `whisper-1`. Thus the transcript may diverge somewhat from the model's interpretation, and should be treated as a rough guide.

## OpenAPI definition

```yaml
openapi: 3.0.0
info:
  title: OpenAI API
  version: 2.3.0
servers:
  - url: https://api.openai.com/v1
components:
  schemas:
    RealtimeServerEventConversationItemInputAudioTranscriptionCompleted:
      type: object
      description: >
        This event is the output of audio transcription for user audio written
        to the 

        user audio buffer. Transcription begins when the input audio buffer is 

        committed by the client or server (in `server_vad` mode). Transcription
        runs 

        asynchronously with Response creation, so this event may come before or
        after 

        the Response events.


        Realtime API models accept audio natively, and thus input transcription
        is a 

        separate process run on a separate ASR (Automatic Speech Recognition)
        model, 

        currently always `whisper-1`. Thus the transcript may diverge somewhat
        from 

        the model's interpretation, and should be treated as a rough guide.
      properties:
        event_id:
          type: string
          description: The unique ID of the server event.
        type:
          type: string
          enum:
            - conversation.item.input_audio_transcription.completed
          description: |
            The event type, must be
            `conversation.item.input_audio_transcription.completed`.
          x-stainless-const: true
        item_id:
          type: string
          description: The ID of the user message item containing the audio.
        content_index:
          type: integer
          description: The index of the content part containing the audio.
        transcript:
          type: string
          description: The transcribed text.
      required:
        - event_id
        - type
        - item_id
        - content_index
        - transcript
      x-oaiMeta:
        name: conversation.item.input_audio_transcription.completed
        group: realtime
        example: |
          {
              "event_id": "event_2122",
              "type": "conversation.item.input_audio_transcription.completed",
              "item_id": "msg_003",
              "content_index": 0,
              "transcript": "Hello, how are you?"
          }
```
