Real-Time Multi-Sign Language Recognition and Translation

Sign language recognition, ISL, ASL, PSL, IPSL, CNN, LSTM, YOLOv8, MediaPipe, real-time translation, bidirectional SLR, LLM, TFLite, gesture recognition, deaf accessibility.

Authors

  • Dr. Tabasum Guledgud Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
  • Mr. Siddarth Hugar Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
  • Ms. Rakshita V Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
  • Ms. Soumya patil Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
  • Ms. Sushmita A Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
  • Ms. Usha A Department of Computer Science and Engineering AGM Rural College of Engineering and Technology,Varur, Hubballi, Karnataka
May 25, 2026
May 27, 2026

Downloads

Sign language constitutes the primary mode of communication for an estimated 466 million deaf and hard-of-hearing individuals worldwide, yet the overwhelming majority of hearing individuals cannot understand it. Existing sign language recognition and translation systems address only a single sign language in isolation, are limited to static gestures or a small lexicon, lack real-time bidirectional translation, and fail to support multiple regional sign languages simultaneously. This paper proposes UnifySign: a unified, AI-powered, real-time framework for multi-sign language recognition and translation, supporting Indian Sign Language (ISL), American Sign Language (ASL), Pakistan Sign Language (PSL), and Indo-Pakistani Sign Language (IPSL). The system integrates YOLOv8 for bounding-box detection, MediaPipe Holistic for 21-point hand landmark extraction, Convolutional Neural Networks (CNN) for static sign classification, Long Short-Term Memory (LSTM) networks for dynamic gesture sequence recognition, and a Large Language Model (Google Flan-T5) for grammatically coherent sentence generation. A bidirectional translation pipeline supports both sign-to-text and speech/text-to-sign directions. Offline mobile deployment (Flutter + TFLite) and a web-based interface (Flask) ensure accessibility across device types. Experimental evaluation on INCLUDE, MS-ASL, WLASL, and a custom PSL/IPSL corpus demonstrates accuracy exceeding 97% for ISL/ASL static recognition, 95% for dynamic gestures, and 91% for sentence-level translation. The proposed system surpasses all ten reviewed state-of-the-art systems in coverage, accuracy, and real-world deployability.