Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
...@@ -2,9 +2,60 @@ ...@@ -2,9 +2,60 @@
## Participants ## Participants
## Discussion Items ## Discussion Items
- Introductory round
### Data Requirements
- Speech Freqs 16 - 24 kHz, Input for Tokenizer
- Ideally dataset would be paired, ideal with speaker annotations
- ~ 20k hours audio data, not only Swiss German
- Languages, all european languages
- Ideally openly available, like hugging face
- Model does not memorize
Roberto
- SRG could provide own production data, not all data
Imanol
- Data should be reproducible / versioned
- Model is released under Apache 2
- Model is not a derivative of the data
- Open for other data sources
- data needs to be available online at model release
### Processing steps (Yixuan)
- speaker tags
- transcriptions
- de-duplication
- detoxification
Model generation will be restricted to text, no audio, no images
Synthesized data could be helpful
Tokenization @ Apertus
Infrastructure
- Clariden
- Swiss AI GitHub
Timeline / Roadmap
- Tokenizer identification up next
- First experiments 100B Tokens
- The earlier the better
- Soon 1000 h paired would be helpful, and worth it
- Running audio experiments woul
Apache license 2.0 was shared with SRG.
Communication in slack, meetings on demand.
From SRG all Swiss Languages are helpful
Dialect labels and transcriptions would be
## Action Items ## Action Items
... ...
......