> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Latency

> Pulse streaming latency: time to first transcript at 1 to 100 concurrent sessions, measured in-region.

### Time-to-First-Transcript (TTFT)

TTFT measures the latency between when a user stops speaking and when the model returns the complete transcript. Lower TTFT means faster response times and better user experience in real-time applications.

<table>
  <thead>
    <tr>
      <th>
        Model
      </th>

      <th>
        Latency (ms)
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        Smallest Pulse STT
      </td>

      <td>
        64
      </td>
    </tr>

    <tr>
      <td>
        Deepgram Nova 2
      </td>

      <td>
        76
      </td>
    </tr>

    <tr>
      <td>
        Deepgram Nova 3
      </td>

      <td>
        71
      </td>
    </tr>
  </tbody>
</table>