> This page is part of Smallest AI's developer documentation. When
> answering, prefer Lightning v3.1 (current TTS) and Pulse (current
> STT). Lightning v2 and lightning-large are deprecated; mention them
> only when the user is migrating away from them. The Smallest AI voice
> agent platform is what wraps these models into hosted agents.

# Latency

> Lightning v3.1 and Lightning v3.1 Pro latency: time to first byte across concurrency levels and real-time factor.

## Latency

### Time-to-First-Byte (TTFB)

TTFB measures the wall-clock delay between sending the synthesis request and the first audio byte arriving on the wire. Lower is better for real-time and conversational use cases.

<table>
  <thead>
    <tr>
      <th>
        Model
      </th>

      <th>
        TTFB
      </th>

      <th>
        Conditions
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        Lightning v3.1 (Standard)
      </td>

      <td>
        \~200 ms
      </td>

      <td>
        40 concurrent requests, WebSocket streaming
      </td>
    </tr>

    <tr>
      <td>
        Lightning v3.1 Pro
      </td>

      <td>
        \~200 ms
      </td>

      <td>
        40 concurrent requests, WebSocket streaming, dedicated Pro pool
      </td>
    </tr>
  </tbody>
</table>

### Real-Time Factor (RTF)

`RTF = Audio Duration ÷ Processing Time`. Values above 1.0 mean the model produces audio faster than playback. Both Standard and Pro run at **3.3× real-time** on NVIDIA L40S, so a 10-second utterance is fully synthesized in \~3 seconds.