01
lite-64kSpeed & volume
Lite
Real-Time FastLow-latency inference for customer-facing chat, high-volume classification, and real-time AI agents.Explore Lite64Ktoken context window
Real workloadResolve a live customer request.Latency first
- Input
- Support history + current message
- Route
- lite-64k
- Output
- Low-latency streamed answer
