AI / Model series
The PICO Series.
A speech-to-speech model. Voice in, voice out, without the detour through text, and one of the more developed ideas in our pipeline.
What's here?
Models in this series.
What the series is for
PICO is built around a single idea: hear speech and answer in speech directly, without converting to text in the middle. That removes the lag and the loss that come from stitching separate speech-to-text and text-to-speech systems together, and makes for conversation that feels immediate.
PICO-1, the first release, is text-to-speech rather than full speech-to-speech: it establishes the streaming, low-latency playback approach the rest of the series builds on.
PICO's pace slowed once we set a clear order of priority across projects; see Evaluating scope for why. Work continues, just more slowly than Deriva and Hikaru.
Related
- PICO-1, the model page
- PICO-1 enters development, the announcement
- Evaluating scope, why PICO slowed down
- All model series, the full lineup
- Our research, the methods behind the models