Changes
Page history
Update Swiss German Speech Data Meeting Notes
authored
Apr 01, 2026
by
Michael Graber
Hide whitespace changes
Inline
Side-by-side
Swiss-German-Speech-Data---Meeting-Notes.md
View page @
2ff1d14d
...
...
@@ -2,6 +2,31 @@
### Participants
Yixuan Xu, Daniel Perruchoud, Michael Graber
### Previous Action Items
*
[ ] Roberto: Licence agreement, also partially for Regionaljournal & Espresso broadcasts (needs to be commercially usable and openly shareable)
*
[x] FHNW : translate SFT datasets
*
[x] Prepare SRF dataset plan, share with Yixuan
### Discussion Points
*
FHNW :
\~
3500h SRF almost ready, sanity checking ongoing,
*
chunk size: 2050 s per batch would be possible, good would be
\~
30 - 50 s chunks
*
global ID for original audio clip + segement should be conserved, if possible speaker ID(s)
*
Yixuan : Zürich Parliament data to be captioned with Qwen-3 omni caption, to be added to training
*
maybe 10% of more complicated audio will be captioned like this, i.e. 50k hours
*
Roberto: first licence agreement draft available from Lars at SRF, will be shared with SVEN
*
Yixuan: audio data with translation to english, could be used for additional (80k hours, https://huggingface.co/datasets/nvidia/Granary)
*
Would it be possible to translate Swiss German audio to english
### New Action Items
*
[ ] FHNW: will deliver the Espresso and Regionaljournal data via huggingface dataset with download token to Yixuan
### Participants
Roberto Salomone, Daniel Perruchoud, Michael Graber
### Previous Action Items
...
...
...
...