Changes
Page history
Update Swiss German Speech Data Meeting Notes
authored
Nov 13, 2025
by
Michael Graber
Show whitespace changes
Inline
Side-by-side
Swiss-German-Speech-Data---Meeting-Notes.md
View page @
6ba2772c
# Meeting, November 13, 2025
#
# Meeting, November 13, 2025
## Participants
### Participants
Yixuan Xu, Seven Najem-Meyer, Imanol Schlag, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
## Discussion Items
##
#
Discussion Items
-
Introductory round
####
Introductory round
### Data Requirements
###
#
Data Requirements
-
Speech Freqs 16 - 24 kHz, Input for Tokenizer
-
Speech Freqs 16 - 24 kHz, Input for Tokenizer
-
Ideally dataset would be paired,
ideal with speaker annotation
s
-
Ideally dataset would be paired,
even more ideal including speaker annotations and dialect label
s
-
~ 20k hours audio data, not only Swiss German
-
Goal:
~ 20k hours audio data, not only Swiss German
-
Languages,
all
e
uropean languages
-
Audio of
all
E
uropean languages
will be included
-
Ideally openly available, like hugging face
-
Ideally openly available, like hugging face
-
Model does not memorize
Roberto
-
SRG could provide own production data, not all data
Imanol
-
Data should be reproducible / versioned
-
Data should be reproducible / versioned
-
Data needs to be available online at model release -> reproducibility
-
SRF quality should be good enough
-
All Swiss Languages from SRG are welcome
Re licencing (stressed by Immanol):
-
The model does not memorize (due to Goldfish loss)
-
Hence the model is not a derivative of the data
-
Due to lack of memorization, license constraints of data does not need to apply for the models
-
Model is released under Apache 2
-
Model is released under Apache 2
-
Model is not a derivative of the data
-
Open for other data sources
-
data needs to be available online at model release
### Processing steps (Yixuan)
#### Possible Sources Roberto
-
Roberto: SRG could provide own production data, not all data
-
Apertus: Open for other data sources
-
Synthesized data could be helpful as well
#### Processing steps
-
speaker tags
-
speaker tags
-
transcriptions
-
transcriptions
-
de-duplication
-
de-duplication
-
detoxification
-
detoxification
-
Tokenization is done by Apertus Team
Model generation will be restricted to text, no audio, no images
#### Infrastructure
-
Clariden can be used
Synthesized data could be helpful
-
Ultimately, codes should be available through Swiss AI GitHub
Tokenization @ Apertus
Infrastructure
-
Clariden
-
Swiss AI GitHub
Timeline / Roadmap
####
Timeline / Roadmap
-
Tokenizer identification up next
-
Tokenizer identification up next
-
First experiments 100B Tokens
-
First experiments 100B Tokens will follow
-
The earlier the better
-
It would be helpful if ~ 1000 h paired Swiss German data would be available
-
Soon 1000 h paired would be helpful, and worth it
-
Running audio experiments woul
Apache license 2.0 was shared with SRG.
Communication in slack, meetings on demand.
From SRG all Swiss Languages are helpful
Dialect labels and transcriptions would be
#### Varia
-
Model generation will be restricted to text, no audio, no images
-
Apache license 2.0 was shared with SRG.
-
Channel on Swiss-AI slack will be created
-
To start, meetings will be held weekly, same timeslot
## Action Items
##
#
Action Items
-
[ ] Complete data source list (@michael.graber, @daniel.perruchoud )
-
[ ] Define roadmap (@michael.graber, @daniel.perruchoud )
-
[ ] Define roadmap (@michael.graber, @daniel.perruchoud )
-
[ ] Identify list of downloadable casts complying with Apertus co (@Roberto & @Vincent / SRF)
-
[ ] Setup meeting series in same slot (@michael.graber)
\ No newline at end of file