Abstract:Efforts to make online media accessible to a regional audience have picked up pace in recent years with multilingual captioning and keyboards. However, techniques to extend this access to people with hearing loss are limited. Further, owing to a lack of structure in the education of hearing impaired and regional differences, the issue of standardization of Indian Sign Language (ISL) has been left unaddressed, forcing educators to rely on the local language to support the ISL structure, thereby creating an array of correlations for each object, hindering the language building skills of a student. This paper aims to present a useful technology that can be used to leverage online resources and make them accessible to the hearing-impaired community in their primary mode of communication. Our tool presents an avenue for the early development of language learning and communication skills essential for the education of children with a profound hearing loss. With the proposed technology, we aim to provide a standardized teaching and learning medium to a classroom setting that can utilize and promote ISL. The goals of our proposed system involve reducing the burden of teachers to act as a valuable teaching aid. The system allows for easy translation of any online video and correlation with ISL captioning using a 3D cartoonish avatar aimed to reinforce classroom concepts during the critical period. First, the video gets converted to text via subtitles and speech processing methods. The generated text is understood through NLP algorithms and then mapped to avatar captions which are then rendered to form a cohesive video alongside the original content. We validated our results through a 6-month period and a consequent 2-month study, where we recorded a 37% and 70% increase in performance of students taught using Sign captioned videos against student taught with English captioned videos. We also recorded a 73.08% increase in vocabulary acquisition through signed aided videos.

Caption positioning structure for hard of hearing people using deep learning method

A Survey Study on Automatic Subtitle Synchronization and Positioning System for Deaf and Hearing Impaired People

Image Recognition Using Text and Audio Translation for the Visually Challenged

Video accessibility enhancement for hearing-impaired users

Addressing visual impairments: Essential software requirements for image caption solutions

See-Through Captions: Real-Time Captioning on Transparent Display for Deaf and Hard-of-Hearing People

Video captioning – a survey

Automated 3D sign language caption generation for video

A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning

Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment

Motion Guided Region Message Passing for Video Captioning

Joint Generation of Captions and Subtitles with Dual Decoding

Towards Real Time Egocentric Segment Captioning for The Blind and Visually Impaired in RGB-D Theatre Images

Automated Audio Captioning with Recurrent Neural Networks

A Lightweight Visual Understanding System for Enhanced Assistance to the Visually Impaired Using an Embedded Platform

Portable Camera-Based Product Label Reading For Blind People

Positional Self-attention Based Hierarchical Image Captioning.

Caption Anything: Interactive Image Description with Diverse Multimodal Controls

D-CNN: A New model for Generating Image Captions with Text Extraction Using Deep Learning for Visually Challenged Individuals

Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention

Delving Deeper into the Decoder for Video Captioning