From Accessibility Needs to AI System Design: A User-Centered Study of Image-To-Speech Assistive Technology for People with Visual Impairments in Thanh Hoa, Vietnam

Author's Information:

Hoang Anh Cong

Hong Duc University, Vietnam

Vol 3 No 9 (2026):Volume 03 Issue 09 September 2026

Page No.: 1150-1158

Abstract:

Advances in artificial intelligence (AI), computer vision, optical character recognition (OCR), and speech synthesis are creating new possibilities for helping people with visual impairments access visual information. However, the effectiveness of assistive technology depends not only on model accuracy but also on its fit with users' needs, digital capabilities, and contexts of use. This study identifies information-access barriers, technology-use needs, and functional requirements for a Vietnamese image-to-speech system designed to support people with visual impairments in Thanh Hoa, Vietnam. A mixed-methods descriptive and applied design was employed, combining questionnaire surveys, direct interviews, field observation, and documented field evidence. A total of 230 valid questionnaires were analyzed, comprising 100 responses from people with visual impairments, 120 from guardians/support persons, and 10 from representatives of supporting agencies and organizations. Among the 100 users with visual impairments, 79% had severe visual impairment or complete blindness. The findings indicate particularly strong needs for reading text from images, describing natural scenes, recognizing obstacles and signs, and receiving results in Vietnamese speech. Smartphones provide a promising deployment platform, but users' operational capabilities are uneven, especially in camera use, touch interaction, and error handling. On this basis, the study proposes an architecture integrating image acquisition and quality checking, OCR/AI-based image description, Vietnamese language processing, text-to-speech synthesis, and an audio-first interface. Cross-cutting requirements include minimal interaction steps, audio-guided camera operation, TalkBack/VoiceOver compatibility, rapid feedback, partial offline functionality, and user-data protection. The study demonstrates that translating context-specific user needs into technical requirements is essential for developing AI assistive technologies that are usable in practice and responsive to local conditions.

KeyWords:

visual impairment, artificial intelligence, assistive technology, image-to-speech, accessibility, user-centered design, Thanh Hoa.

References:

  1. Apple. (n.d.). Turn on and practice VoiceOver on iPhone. Apple Support. https://support.apple.com/guide/iphone/turn-on-and-practice-voiceover-iph3e2e415f/ios
  2. Be My Eyes. (n.d.). Be My Eyes app: Bringing sight to blind and low-vision people. https://www.bemyeyes.com/
  3. Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., & Zitnick, C. L. (2015). Microsoft COCO captions: Data collection and evaluation server. arXiv:1504.00325. https://arxiv.org/abs/1504.00325
  4. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723
  5. Google. (n.d.). Get started on Android with TalkBack. Android Accessibility Help. https://support.google.com/accessibility/android/answer/6283677
  6. Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., & Bigham, J. P. (2020). VizWiz: Nearly real-time answers to visual questions. International Journal of Computer Vision, 128, 1-20. https://doi.org/10.1007/s11263-019-01266-6
  7. Herdade, S., Kappeler, A., Boakye, K., & Soares, J. (2019). Image captioning: Transforming objects into words. Advances in Neural Information Processing Systems, 32, 11137-11147.
  8. Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3128-3137. https://doi.org/10.1109/CVPR.2015.7298932
  9. Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In Computer Vision - ECCV 2014 (pp. 740-755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
  10. Microsoft. (2023, December 4). Seeing AI app launches on Android: Including new and updated features and new languages. Microsoft Accessibility Blog.
  11. Phuong, N. T. (2026). Digital preservation of Vietnamese cultural heritage: Opportunities and limitations in the age of smart technology. Preservation, Digital Technology & Culture, 55(2), 117-129. https://doi.org/10.1515/pdtc-2025-0035
  12. Phuong, N. T., & Lam, N. T. N. (2025). Evaluating the effectiveness of cultural heritage communication based on local community feedback: A case study of Hanoi city. Scientific Culture, 11(3), 1-11. https://doi.org/10.5281/zenodo.17379320
  13. Smith, R. (2007). An overview of the Tesseract OCR engine. Ninth International Conference on Document Analysis and Recognition, 629-633. https://doi.org/10.1109/ICDAR.2007.4376991
  14. United Nations. (2006). Convention on the Rights of Persons with Disabilities. https://www.un.org/development/desa/disabilities/convention-on-the-rights-of-persons-with-disabilities.html
  15. Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3156-3164. https://doi.org/10.1109/CVPR.2015.7298935
  16. W3C. (2023). Web Content Accessibility Guidelines (WCAG) 2.2. World Wide Web Consortium. https://www.w3.org/TR/WCAG22/
  17. Wang, Y., Skerry-Ryan, R. J., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., & Saurous, R. A. (2017). Tacotron: Towards end-to-end speech synthesis. Interspeech 2017, 4006-4010. https://doi.org/10.21437/Interspeech.2017-1452
  18. Whang, S. E., Roh, Y., Song, H., & Lee, J.-G. (2023). Data collection and quality challenges in deep learning: A data-centric AI perspective. The VLDB Journal, 32, 791-813. https://doi.org/10.1007/s00778-022-00775-9
  19. World Health Organization. (2019). World report on vision. World Health Organization. https://www.who.int/publications/i/item/9789241516570
  20. World Health Organization. (2023). Disability and health. World Health Organization. https://www.who.int/news-room/fact-sheets/detail/disability-and-health