Changes
Page history
Update Swiss German Speech Data Meeting Notes
authored
Nov 13, 2025
by
Michael Graber
Show whitespace changes
Inline
Side-by-side
Swiss-German-Speech-Data---Meeting-Notes.md
View page @
51ba3263
...
@@ -2,9 +2,60 @@
...
@@ -2,9 +2,60 @@
## Participants
## Participants
## Discussion Items
## Discussion Items
-
Introductory round
### Data Requirements
-
Speech Freqs 16 - 24 kHz, Input for Tokenizer
-
Ideally dataset would be paired, ideal with speaker annotations
-
~ 20k hours audio data, not only Swiss German
-
Languages, all european languages
-
Ideally openly available, like hugging face
-
Model does not memorize
Roberto
-
SRG could provide own production data, not all data
Imanol
-
Data should be reproducible / versioned
-
Model is released under Apache 2
-
Model is not a derivative of the data
-
Open for other data sources
-
data needs to be available online at model release
### Processing steps (Yixuan)
-
speaker tags
-
transcriptions
-
de-duplication
-
detoxification
Model generation will be restricted to text, no audio, no images
Synthesized data could be helpful
Tokenization @ Apertus
Infrastructure
-
Clariden
-
Swiss AI GitHub
Timeline / Roadmap
-
Tokenizer identification up next
-
First experiments 100B Tokens
-
The earlier the better
-
Soon 1000 h paired would be helpful, and worth it
-
Running audio experiments woul
Apache license 2.0 was shared with SRG.
Communication in slack, meetings on demand.
From SRG all Swiss Languages are helpful
Dialect labels and transcriptions would be
## Action Items
## Action Items
...
...
...
...