- [ ] @Roberto: Confirm huggingface dump for 100 Sekunden Wissen, Licence for hugging face cc by 4.0 nc (?)
- [ ] @Roberto: Confirm huggingface dump for 100 Sekunden Wissen, Licence for hugging face cc by 4.0 nc (?)
- [ ] @Vincent & @FHNW : create transcriptions
- [ ] @Vincent & @fhnw : create transcriptions
## Meeting, November 20, 2025
## Meeting, November 20, 2025
### Participants
### Participants
Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Previous Action Items
### Previous Action Items
...
@@ -70,15 +89,15 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
...
@@ -70,15 +89,15 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
### Discussion Items
### Discussion Items
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (~ 100 TB of data)
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (\~ 100 TB of data)
- it's currently on S3-buckets on AWS, SSR will provide list of pointers
- it's currently on S3-buckets on AWS, SSR will provide list of pointers
- long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team)
- long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team)
- Datasets: ~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Datasets: \~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Contribution Modalities
- Contribution Modalities
- interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training
- interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training