Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
# Meeting, November 13, 2025
## Meeting, November 13, 2025
## Participants
### Participants
Yixuan Xu, Seven Najem-Meyer, Imanol Schlag, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
## Discussion Items
### Discussion Items
- Introductory round
#### Introductory round
### Data Requirements
#### Data Requirements
- Speech Freqs 16 - 24 kHz, Input for Tokenizer
- Ideally dataset would be paired, ideal with speaker annotations
- ~ 20k hours audio data, not only Swiss German
- Languages, all european languages
- Ideally dataset would be paired, even more ideal including speaker annotations and dialect labels
- Goal: ~ 20k hours audio data, not only Swiss German
- Audio of all European languages will be included
- Ideally openly available, like hugging face
- Model does not memorize
Roberto
- SRG could provide own production data, not all data
Imanol
- Data should be reproducible / versioned
- Data needs to be available online at model release -> reproducibility
- SRF quality should be good enough
- All Swiss Languages from SRG are welcome
Re licencing (stressed by Immanol):
- The model does not memorize (due to Goldfish loss)
- Hence the model is not a derivative of the data
- Due to lack of memorization, license constraints of data does not need to apply for the models
- Model is released under Apache 2
- Model is not a derivative of the data
- Open for other data sources
- data needs to be available online at model release
### Processing steps (Yixuan)
#### Possible Sources Roberto
- Roberto: SRG could provide own production data, not all data
- Apertus: Open for other data sources
- Synthesized data could be helpful as well
#### Processing steps
- speaker tags
- transcriptions
- de-duplication
- detoxification
- Tokenization is done by Apertus Team
Model generation will be restricted to text, no audio, no images
Synthesized data could be helpful
Tokenization @ Apertus
Infrastructure
- Clariden
- Swiss AI GitHub
#### Infrastructure
- Clariden can be used
- Ultimately, codes should be available through Swiss AI GitHub
Timeline / Roadmap
#### Timeline / Roadmap
- Tokenizer identification up next
- First experiments 100B Tokens
- The earlier the better
- Soon 1000 h paired would be helpful, and worth it
- Running audio experiments woul
Apache license 2.0 was shared with SRG.
Communication in slack, meetings on demand.
From SRG all Swiss Languages are helpful
Dialect labels and transcriptions would be
- First experiments 100B Tokens will follow
- It would be helpful if ~ 1000 h paired Swiss German data would be available
#### Varia
- Model generation will be restricted to text, no audio, no images
- Apache license 2.0 was shared with SRG.
- Channel on Swiss-AI slack will be created
- To start, meetings will be held weekly, same timeslot
## Action Items
### Action Items
- [ ] Complete data source list (@michael.graber, @daniel.perruchoud )
- [ ] Define roadmap (@michael.graber, @daniel.perruchoud )
- [ ] Identify list of downloadable casts complying with Apertus co (@Roberto & @Vincent / SRF)
- [ ] Setup meeting series in same slot (@michael.graber)
\ No newline at end of file