Captions provide the spoken words, but in a crowded room, it is crucial to identify the speaker. CaptiVision (www.captivision.app) is an accessibility app that helps hard-of-hearing professionals follow group conversations by combining live, speaker-labeled captions with active speaker tracking and multi-camera video feeds. CaptiVision provides a unified visual stream that eliminates split focus for Deaf and Hard of Hearing individuals in classrooms and meeting spaces.
This past week was International Week of the Deaf, and to mark the occasion, I’m excited to launch CaptiVision. This is the story of how it all happened.
In 2025, I was selected for the Mandela Washington Fellowship, the flagship program of the U.S. Government’s Young African Leaders Initiative (YALI). I was hosted at the University of Texas at Austin alongside a cohort of accomplished leaders. That is where I met Mamy Sira, a deaf business owner and disability advocate from The Gambia. Attending classes directly behind Mamy Sira, I learned about the challenges Deaf and Hard of Hearing individuals encounter when attending regular lectures and meetings.
The University of Texas at Austin provided extensive accommodations for the program, including American Sign Language (ASL) interpreters, desktop microphones, and real-time captioning software. However, during intensive lectures, I noticed a persistent struggle. To follow discussions, Mamy Sira had to repeatedly shift her focus between the lecture slides, the live ASL interpreter, the laptop screen displaying captions, and whichever participant was speaking across the room. Tracking dynamic classroom dialogue required continuous visual scanning, which resulted in unnecessary fatigue and neck pain during an already demanding six-week course.
I envisioned a system capable of integrating all visual cues. I discussed the concept with Mamy Sira and ASL interpreters Aisha S. Terp and Dr. Saint Hardy, subsequently obtaining permission from Prof. Rodney Northern to conduct a live trial. Utilizing borrowed equipment from the McCombs School of Business AV support desk, I assembled a rudimentary prototype employing a laptop, an iPhone, a USB webcam, and Zoom. The initial version was rudimentary: the quality was subpar, the captioning software and Zoom were separate applications, and navigating multiple Zoom windows to locate the speaker remained challenging, yet the fundamental function was evident. Aisha remarked that it was how every class session should have been delivered from the start.
Upon returning to Ethiopia, I embarked on the development of a standalone platform named CaptiVision. Translating the prototype into a dependable application necessitated overcoming intricate engineering challenges: integrating computer vision, machine learning, active-speaker detection, peer-to-peer networking, facial recognition, on-device captioning, speaker diarization, and multi-stream video and audio processing.
After months of development and overcoming several points of near-burnout, I completed the initial build. CaptiVision delivers a unified visual stream designed to eliminate split focus for Deaf and Hard of Hearing individuals in classrooms and meeting spaces, with ongoing efforts underway to make the app entirely free for K-12 students worldwide.
Spare Mic - Wireless Microphone
This past week was International Week of the Deaf, and to mark the occasion, I’m excited to launch CaptiVision. This is the story of how it all happened.
In 2025, I was selected for the Mandela Washington Fellowship, the flagship program of the U.S. Government’s Young African Leaders Initiative (YALI). I was hosted at the University of Texas at Austin alongside a cohort of accomplished leaders. That is where I met Mamy Sira, a deaf business owner and disability advocate from The Gambia. Attending classes directly behind Mamy Sira, I learned about the challenges Deaf and Hard of Hearing individuals encounter when attending regular lectures and meetings.
The University of Texas at Austin provided extensive accommodations for the program, including American Sign Language (ASL) interpreters, desktop microphones, and real-time captioning software. However, during intensive lectures, I noticed a persistent struggle. To follow discussions, Mamy Sira had to repeatedly shift her focus between the lecture slides, the live ASL interpreter, the laptop screen displaying captions, and whichever participant was speaking across the room. Tracking dynamic classroom dialogue required continuous visual scanning, which resulted in unnecessary fatigue and neck pain during an already demanding six-week course.
I envisioned a system capable of integrating all visual cues. I discussed the concept with Mamy Sira and ASL interpreters Aisha S. Terp and Dr. Saint Hardy, subsequently obtaining permission from Prof. Rodney Northern to conduct a live trial. Utilizing borrowed equipment from the McCombs School of Business AV support desk, I assembled a rudimentary prototype employing a laptop, an iPhone, a USB webcam, and Zoom. The initial version was rudimentary: the quality was subpar, the captioning software and Zoom were separate applications, and navigating multiple Zoom windows to locate the speaker remained challenging, yet the fundamental function was evident. Aisha remarked that it was how every class session should have been delivered from the start.
Upon returning to Ethiopia, I embarked on the development of a standalone platform named CaptiVision. Translating the prototype into a dependable application necessitated overcoming intricate engineering challenges: integrating computer vision, machine learning, active-speaker detection, peer-to-peer networking, facial recognition, on-device captioning, speaker diarization, and multi-stream video and audio processing.
After months of development and overcoming several points of near-burnout, I completed the initial build. CaptiVision delivers a unified visual stream designed to eliminate split focus for Deaf and Hard of Hearing individuals in classrooms and meeting spaces, with ongoing efforts underway to make the app entirely free for K-12 students worldwide.