research pipeline · 2026
Moshi turn-taking data pipeline
Reproducible data pipeline for preparing conversational audio and turn-control labels for dialogue-model adaptation.
This pipeline supported my MSc dissertation on adapting full-duplex dialogue models. These models require more than transcripts: training data must preserve timing, speaker state, audio-codec tokens and labels describing future user activity and turn control.
The pipeline discovers suitable conversational recordings, performs speaker diarisation and transcription, encodes audio and generates aligned auxiliary labels. Explicit stages and recorded provenance made model conditions easier to reproduce and compare.