Update Swiss German Speech Data Meeting Notes authored by Michael Graber's avatar Michael Graber
...@@ -7,23 +7,34 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber ...@@ -7,23 +7,34 @@ Yixuan Xu, Daniel Perruchoud, Michael Graber
### Previous Action Items ### Previous Action Items
* [ ] Roberto: Licence agreement, also partially for Regionaljournal & Espresso broadcasts (needs to be commercially usable and openly shareable) * [ ] Roberto: Licence agreement, also partially for Regionaljournal & Espresso broadcasts (needs to be commercially usable and openly shareable)
* [ ] FHNW : translate SFT datasets * [ ] FHNW : translate SFT datasets
### Discussion Points ### Discussion Points
* Transcription Evaluation for SRF broadcast transcripts
* FHNW models perform substantially better than current SRF solution
* ![TranscriptionRatingDistributionSRFvsFHNW.png](uploads/f4f0cc084aa9daae243cf90cb83e3485/TranscriptionRatingDistributionSRFvsFHNW.png){width="504" height="360"}
* ![TranscriptionModelRatings_FHNW-1vsSRF.png](uploads/1a60f756ecd7ce368cebef5a11aabeb0/TranscriptionModelRatings_FHNW-1vsSRF.png){width="461" height="346"}
* ![TranscriptionModelRatings_FHNW-2vsSRF.png](uploads/fad633a86e28d58200bb69dfaa4204e5/TranscriptionModelRatings_FHNW-2vsSRF.png){width="461" height="346"}
* Rating scale:
1 - UNUSABLE: Largely incomprehensible; re-transcription would be faster than correction.
2 - POOR: General content recognizable, but frequent errors; substantial post-editing required.
3 - ADEQUATE: Essential content correct, errors in difficult passages; moderate post-editing needed.
4 - GOOD: High accuracy, errors only with rare words or poor audio quality; minor corrections needed.
5 - EXCELLENT: Nearly error-free, even with technical vocabulary and difficult conditions; practically publication-ready.
* \-\> will use FHNW models to generate transcripts, as soon as licencing situation is clarified
* SFT datasets in discussion for Swiss German translation (text and audio) * SFT datasets in discussion for Swiss German translation (text and audio)
* https://huggingface.co/datasets/deepset/germanquad * https://huggingface.co/datasets/deepset/germanquad
https://huggingface.co/datasets/gpt-omni/VoiceAssistant-400K/viewer/default/train?row=27 https://huggingface.co/datasets/gpt-omni/VoiceAssistant-400K/viewer/default/train?row=27
https://huggingface.co/datasets/lawinstruct/lawinstruct (have to extract german (de)) https://huggingface.co/datasets/lawinstruct/lawinstruct (have to extract german (de))
* FHNW has to identify suitable models (not trained on proprietary data) for translation * FHNW has to identify suitable models (not trained on proprietary data) for translation
* Transcription Evaluation for SRF broadcast transcripts
* FHNW models perform substantially better than current SRF solution
* ![TranscriptionRatingDistributionSRFvsFHNW.png](uploads/f4f0cc084aa9daae243cf90cb83e3485/TranscriptionRatingDistributionSRFvsFHNW.png){width=504 height=360}
* ![TranscriptionModelRatings_FHNW-1vsSRF.png](uploads/1a60f756ecd7ce368cebef5a11aabeb0/TranscriptionModelRatings_FHNW-1vsSRF.png){width=461 height=346}
* ![TranscriptionModelRatings_FHNW-2vsSRF.png](uploads/fad633a86e28d58200bb69dfaa4204e5/TranscriptionModelRatings_FHNW-2vsSRF.png){width=461 height=346}
* \-\> will use FHNW models to generate transcripts, as soon as licencing situation is clarified
### New Action Items ### New Action Items
... ...
......