FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator

Massa Baali, Ahmed Ali

Research output: Contribution to journalConference articlepeer-review

1 Citation (Scopus)

Abstract

This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelization, followed by the commentator's speech synthesizer. Our pipeline included Arabic automatic speech recognition for data labeling, CTC segmentation, transcription vowelization to match speech, and fine-tuning the TTS. Our system is capable of generating speech with its acoustic environment within limited 15 minutes of football commentator recording. Our prototype is generalizable and can be easily applied to different domains and languages.

Original languageEnglish
Pages (from-to)5249-5250
Number of pages2
JournalProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Volume2023-August
Publication statusPublished - 2023
Event24th International Speech Communication Association, Interspeech 2023 - Dublin, Ireland
Duration: 20 Aug 202324 Aug 2023

Keywords

  • speech recognition
  • text-to-speech

Fingerprint

Dive into the research topics of 'FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator'. Together they form a unique fingerprint.

Cite this