Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
......@@ -38,12 +38,12 @@ Re licencing (stressed by Immanol):
#### Infrastructure
- Clariden can be used
- Ultimately, codes should be available through Swiss AI GitHub
- Ultimately, codes should be available through Swiss AI GitHub, now FHNW GitLab is ok
#### Timeline / Roadmap
- Tokenizer identification up next
- First experiments 100B Tokens will follow
- It would be helpful if ~ 1000 h paired Swiss German data would be available
- It would be helpful if ~ 1000 h paired Swiss German data would be available soon
#### Varia
- Model generation will be restricted to text, no audio, no images
......
......