Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
...@@ -15,7 +15,7 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber ...@@ -15,7 +15,7 @@ Yixuan Xu, Seven Najem-Meyer, Vincent Demotz, Daniel Perruchoud, Michael Graber
- SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (~ 100 TB of data) - SSR Contribution: list has been identified, contains own production / publicly available data, sharing modalities have to be worked-out, contract details are being worked out, research purpose use will be ok (~ 100 TB of data)
- it's currently on S3-buckets on AWS, SSR will provide list of pointers - it's currently on S3-buckets on AWS, SSR will provide list of pointers
- long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team) - long-term storage remains to be solved, huggingface hosting would be good (Sven, Apertus Team)
- Datasets: ~ 1300 h Swiss German Parliament data (transcribed) right away - Datasets: ~ 1300 h Swiss German Parliament data (transcribed) can be contributed right away
- Contribution Modalities - Contribution Modalities
- interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training - interleaving data preparation to be done (e.g. 30 tokens speech, 30 tokens text) for pre-training
- timestamps are helpful - timestamps are helpful
... ...
......