Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
## Meeting, December 17, 2025
### Previous Action Items
- [ ] @Yixuan : propose interleaving pipeline meeting dates once he's back from travelling
- [ ] @Roberto: Confirm huggingface dump for 100 Sekunden Wissen, Licence for hugging face cc by 4.0 nc (?)
- [x] @Vincent & @fhnw : create transcriptions
- [x] Detailing Processing Steps for SSR data in next meeting (@vincent, @michael, @daniel)
- [x] @sven : 1-pager for SRG SSR
- [x] @Roberto: share 25 broadcast mp5 via Google Drive
### Discussion Items
* Planning interleaving encoding meeting
* SSR licence status
### New Action Items
## Meeting, December 10, 2025
......@@ -6,13 +23,14 @@ Meeting cancelled. Next meeting planned for December 17.
## Meeting, December 3, 2025
### Participants
Yixuan Xu, Seven Najem-Meyer, Roberto Salomone, Michael Graber
### Previous Action Items
- [ ] @Yixuan : propose interleaving pipeline meeting dates once he's back from travelling
- [ ] @Roberto: Confirm huggingface dump for 100 Sekunden Wissen, Licence for hugging face cc by 4.0 nc (?)
- [ ] @Vincent & @FHNW : create transcriptions
- [ ] @Vincent & @fhnw : create transcriptions
- [ ] Detailing Processing Steps for SSR data in next meeting (@vincent, @michael, @daniel)
### Discussion Items
......@@ -23,13 +41,14 @@ Yixuan Xu, Seven Najem-Meyer, Roberto Salomone, Michael Graber
- Next weekend will be skipped due to unavailability of multiple participants
### New Action Items
- [ ] @sven : 1-pager for SRG SSR
- [ ] @Roberto: share 25 broadcast mp5 via Google Drive
## Meeting, November 27, 2025
### Participants
Yixuan Xu, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Previous Action Items
......@@ -45,20 +64,20 @@ Yixuan Xu, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
- What processing steps could we contribute further?
- Producing an interleaved version of this dataset would be helpful
- SSR data situation, processing steps long-term storage solution
- 100 Sekunden Wissen could be used right away (~ 4000 episodes)
- 100 Sekunden Wissen could be used right away (\~ 4000 episodes)
- Putting it on Huggingface seems ok but will be confirmed by @Roberto
- Transcription of ~50 episodes will be done by FHNW and SSR
- Transcription of \~50 episodes will be done by FHNW and SSR
### New Action Items
- [ ] @Yixuan : propose interleaving pipeline meeting dates
- [ ] @Roberto: Confirm huggingface dump for 100 Sekunden Wissen, Licence for hugging face cc by 4.0 nc (?)
- [ ] @Vincent & @FHNW : create transcriptions
- [ ] @Vincent & @fhnw : create transcriptions
## Meeting, November 20, 2025
### Participants
Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Previous Action Items
......@@ -70,15 +89,15 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Discussion Items
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (~ 100 TB of data)
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (\~ 100 TB of data)
- it's currently on S3-buckets on AWS, SSR will provide list of pointers
- long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team)
- Datasets: ~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Datasets: \~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Contribution Modalities
- interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training
- timestamps are helpful
- [preprocessing slides](https://docs.google.com/presentation/d/1Ni9PF6rG6XVnhnDdPFAXm-cOnJxoIe6n/edit?slide=id.p1#slide=id.p1)
- ~ 6 users on FHNW-side via [form here](https://docs.google.com/forms/d/e/1FAIpQLSdB9I0Y-a18SKcBx7l3veH51WTImUK1nfej8gP6xaqFn4_ZIw/viewform)
- \~ 6 users on FHNW-side via [form here](https://docs.google.com/forms/d/e/1FAIpQLSdB9I0Y-a18SKcBx7l3veH51WTImUK1nfej8gP6xaqFn4_ZIw/viewform)
- users will be added to slack channel
- FHNW Team: Jonas Grüter + Student Team
......@@ -89,10 +108,10 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
- [ ] Onboard student team @ FHNW (@michael.graber, @daniel.perruchoud )
- [ ] Long-term storage solution for SSR data (@sven)
## Meeting, November 13, 2025
### Participants
Yixuan Xu, Seven Najem-Meyer, Imanol Schlag, Roberto Salomone, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Discussion Items
......@@ -100,17 +119,19 @@ Yixuan Xu, Seven Najem-Meyer, Imanol Schlag, Roberto Salomone, Vincent Demotz, D
#### Introductory round
#### Data Requirements
- Speech Freqs 16 - 24 kHz, Input for Tokenizer
- Ideally dataset would be paired, even more ideal including speaker annotations and dialect labels
- Goal: ~ 20k hours audio data, not only Swiss German
- Goal: \~ 20k hours audio data, not only Swiss German
- Audio of all European languages will be included
- Ideally openly available, like hugging face
- Data should be reproducible / versioned
- Data needs to be available online at model release -> reproducibility
- Data needs to be available online at model release -\> reproducibility
- SRF quality should be good enough
- All Swiss Languages from SRG are welcome
Re licencing (stressed by Immanol):
- The model does not memorize (due to Goldfish loss)
- Hence the model is not a derivative of the data
- Due to lack of memorization, license constraints of data do not need to apply for the models
......@@ -118,11 +139,13 @@ Re licencing (stressed by Immanol):
- Model generation will be restricted to text, no audio, no images
#### Possible Sources Roberto
- Roberto: SRG could provide own production data, not all data
- Apertus: Open for other data sources
- Synthesized data could be helpful as well
#### Processing steps
- speaker tags
- transcriptions
- de-duplication
......@@ -130,15 +153,18 @@ Re licencing (stressed by Immanol):
- Tokenization is done by Apertus Team
#### Infrastructure
- Clariden can be used
- Ultimately, codes should be available through Swiss AI GitHub, now FHNW GitLab is ok
#### Timeline / Roadmap
- Tokenizer identification up next
- First experiments with 100B tokens will follow
- It would be helpful if ~ 1000 h paired Swiss German data would be available soon
- It would be helpful if \~ 1000 h paired Swiss German data would be available soon
#### Varia
- Apache license 2.0 was shared with SRG.
- Channel on Swiss-AI slack will be created
- To start, meetings will be held weekly, same timeslot
......
......