Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
# Meeting, November 13, 2025 ## Meeting, November 13, 2025
## Participants ### Participants
Yixuan Xu, Seven Najem-Meyer, Imanol Schlag, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
## Discussion Items ### Discussion Items
- Introductory round #### Introductory round
### Data Requirements #### Data Requirements
- Speech Freqs 16 - 24 kHz, Input for Tokenizer - Speech Freqs 16 - 24 kHz, Input for Tokenizer
- Ideally dataset would be paired, ideal with speaker annotations - Ideally dataset would be paired, even more ideal including speaker annotations and dialect labels
- ~ 20k hours audio data, not only Swiss German - Goal: ~ 20k hours audio data, not only Swiss German
- Languages, all european languages - Audio of all European languages will be included
- Ideally openly available, like hugging face - Ideally openly available, like hugging face
- Model does not memorize
Roberto
- SRG could provide own production data, not all data
Imanol
- Data should be reproducible / versioned - Data should be reproducible / versioned
- Data needs to be available online at model release -> reproducibility
- SRF quality should be good enough
- All Swiss Languages from SRG are welcome
Re licencing (stressed by Immanol):
- The model does not memorize (due to Goldfish loss)
- Hence the model is not a derivative of the data
- Due to lack of memorization, license constraints of data does not need to apply for the models
- Model is released under Apache 2 - Model is released under Apache 2
- Model is not a derivative of the data
- Open for other data sources
- data needs to be available online at model release
### Processing steps (Yixuan) #### Possible Sources Roberto
- Roberto: SRG could provide own production data, not all data
- Apertus: Open for other data sources
- Synthesized data could be helpful as well
#### Processing steps
- speaker tags - speaker tags
- transcriptions - transcriptions
- de-duplication - de-duplication
- detoxification - detoxification
- Tokenization is done by Apertus Team
Model generation will be restricted to text, no audio, no images #### Infrastructure
- Clariden can be used
Synthesized data could be helpful - Ultimately, codes should be available through Swiss AI GitHub
Tokenization @ Apertus
Infrastructure
- Clariden
- Swiss AI GitHub
Timeline / Roadmap #### Timeline / Roadmap
- Tokenizer identification up next - Tokenizer identification up next
- First experiments 100B Tokens - First experiments 100B Tokens will follow
- The earlier the better - It would be helpful if ~ 1000 h paired Swiss German data would be available
- Soon 1000 h paired would be helpful, and worth it
- Running audio experiments woul
Apache license 2.0 was shared with SRG.
Communication in slack, meetings on demand.
From SRG all Swiss Languages are helpful
Dialect labels and transcriptions would be
#### Varia
- Model generation will be restricted to text, no audio, no images
- Apache license 2.0 was shared with SRG.
- Channel on Swiss-AI slack will be created
- To start, meetings will be held weekly, same timeslot
## Action Items ### Action Items
- [ ] Complete data source list (@michael.graber, @daniel.perruchoud )
- [ ] Define roadmap (@michael.graber, @daniel.perruchoud ) - [ ] Define roadmap (@michael.graber, @daniel.perruchoud )
- [ ] Identify list of downloadable casts complying with Apertus co (@Roberto & @Vincent / SRF)
- [ ] Setup meeting series in same slot (@michael.graber)
\ No newline at end of file