@@ -15,7 +15,7 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (~ 100 TB of data)
- it's currently on S3-buckets on AWS, SSR will provide list of pointers
- long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team)
- Datasets: ~ 1300 h Swiss German Parliament data (transcribed) right away
- Datasets: ~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Contribution Modalities
- interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training