Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
......@@ -22,9 +22,13 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber
* Synthetic Swiss German data?
* Daniel: Idiotikon data might be an option
* Permissive licence
* Ideally for stage 2 would be daily conversations, no bias towards political parliament data
* SRF -\> Ablation study could be done (\~ 2000 h)
* SwissGPC? YouTube has two licences, fully permissive and fully restricted, lincences filtering could be done
* Ideally for stage 2 would be daily conversations, no bias towards political parliament data
* SRF -\> Ablation study could be done (\~ 2000 h)
* SwissGPC? YouTube has two licences, fully permissive and fully restricted, lincences filtering could be done
* Current timeline
* End of phase 1, will probably be redone
* Early ablations of phase 2 done
* Focus for now on stage 2 audio data, also TTS
* FHNW: Transcription Evaluation for SRF broadcast transcripts
* FHNW models perform substantially better than current SRF solution
* ![TranscriptionRatingDistributionSRFvsFHNW.png](uploads/f4f0cc084aa9daae243cf90cb83e3485/TranscriptionRatingDistributionSRFvsFHNW.png){width="504" height="360"}
......
......