DOCUMENT-TO-SPEECH SYSTEM WITH MULTILINGUAL SUPPORT AND ADAPTIVE VOICE FEATURESPresented by : Junalyn C. Burgos & Roselyn V. Cela
INTRODUCTION Digital documents rapidly increasing worldwide Manual reading slow and tiring Users need accessible reading solutions Text-to-speech converts text into speech Supports learning and accessibility (Fitria, 2022) Multilingual speech technologies improving (Saeki et al., 2024)
BACKGROUND OF THE STUDY Eye strain from prolonged screen reading Reading fatigue affects many users Visually impaired face accessibility challenges Language barriers limit document understanding Multilingual speech synthesis improving accessibility Need adaptive automated reading system (Isewon et al., 2014; Markopoulos et al., 2023)
General Problem: • Manual reading inefficient, tiring • Limited accessibility for users • Productivity affected by document overload STATEMENT OF THE PROBLEM
Inefficient multi-document reading Eye strain from long reading Limited multilingual support Lack of adaptive voice options No integrated document-to-speech systemSPECIFIC PROBLEMS
General Objective: •Develop system converting documents to speech OBJECTIVES OF THE STUDY
Analyze user and system requirements Design architecture for multiple formats Develop modules: upload, translate, TTS, audio Test and evaluate system functionality Deploy or propose system implementationSPECIFIC OBJECTIVES
CONCEPTUAL FRAMEWORK (IPO MODEL)•Document upload (PDF, Word, TXT, PPTX) •User login credentials •Language selection •Voice style preferencesINPUT•User authentication •Document validation & extraction •Translation if needed •Adaptive TTS conversion Audio playback with karaoke highlighting•Audio playback in selected language •Highlighted text synced with audio •Downloadable audio fileOUTPUTPROCESS
REFERENCES Saeki, T., Wang, G., Morioka, N., Elias, I., Kastner, K., Biadsy, F., Rosenberg, A., Ramabhadran, B., Zen, H., Beaufays, F., & Shemtov, H. (2024). Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data. http://arxiv.org/abs/2402.18932 Isewon, I., Oyelade, J., & Oladipupo, O. (2014). ISSN : 2249-0868 Foundation of Computer Science FCS. In International Journal of Applied Information Systems (IJAIS) (Vol. 7, Issue 2). www.ijais.org Fitria, T. N. (2022). Utilizing Text-to-Speech Technology: Natural Reader in Teaching Pronunciation. JETLEE : Journal of English Language Teaching, Linguistics, and Literature, 2(2), 70–78. https://doi.org/10.47766/jetlee.v2i2.312 Markopoulos, K., Maniati, G., Vamvoukakis, G., Ellinas, N., Vardaxoglou, G., Kakoulidis, P., Oh, J., Jho, G., Hwang, I., Chalamandaris, A., Tsiakoulis, P., & Raptis, S. (2023). Generating Multilingual Gender-Ambiguous Text-to- Speech Voices. http://arxiv.org/abs/2211.00375
Presented by : Junalyn C. Burgos & Roselyn V. CelaThank You