Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
......@@ -13,7 +13,19 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber
* Yixuan shares https://swiss-ai.github.io/benchmark-audio-tokenizer/
* Checkpoints of ablations
* Transcription Evaluation for SRF broadcast transcripts
* SPC_R seems to only contain 482 h after VAD
* TODO at FHNW -\> clarify SPC_R size
* SPC_R has a great influence on performance for cross-modality alignment, but it's too little paired data
* stage 2 training: 30% transcription, 70% continuation
* possible solutions
* transcripts of Gemeinderat Zürich? -\> not available
* Synthetic Swiss German data?
* Daniel: Idiotikon data might be an option
* Permissive licence
* SRF -\> Ablation study could be done (\~ 2000 h)
* FHNW: Transcription Evaluation for SRF broadcast transcripts
* FHNW models perform substantially better than current SRF solution
* ![TranscriptionRatingDistributionSRFvsFHNW.png](uploads/f4f0cc084aa9daae243cf90cb83e3485/TranscriptionRatingDistributionSRFvsFHNW.png){width="504" height="360"}
* ![TranscriptionModelRatings_FHNW-1vsSRF.png](uploads/1a60f756ecd7ce368cebef5a11aabeb0/TranscriptionModelRatings_FHNW-1vsSRF.png){width="461" height="346"}
......@@ -40,6 +52,8 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber
### New Action Items
* Clarify size of SPC_R dataset
## Meeting, March 4, 2026
### Participants
......
......