Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
...@@ -13,7 +13,7 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber ...@@ -13,7 +13,7 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber
* Yixuan shares https://swiss-ai.github.io/benchmark-audio-tokenizer/ * Yixuan shares https://swiss-ai.github.io/benchmark-audio-tokenizer/
* Checkpoints of ablations * Checkpoints of ablations
* SPC_R seems to only contain 482 h after VAD * SPC_R seems to only contain 482 h after VAD, 524 h before
* TODO at FHNW -\> clarify SPC_R size * TODO at FHNW -\> clarify SPC_R size
* SPC_R has a great influence on performance for cross-modality alignment, but it's too little paired data * SPC_R has a great influence on performance for cross-modality alignment, but it's too little paired data
* stage 2 training: 30% transcription, 70% continuation * stage 2 training: 30% transcription, 70% continuation
... ...
......