Hi all. First post here, so a quick intro: I'm Jose, I've spent 29 years working in IT and cybersecurity, and I'm sighted, so my main job in this thread is to listen. I'm not selling anything. I have an idea for a free, open source tool, and before I write a single line of code I want to hear from the people who would actually use it whether it's worth building, and what would make it trustworthy. If the honest answer is "we don't need this," that is a genuinely useful answer and I want to hear it.
The idea in plain terms: a browser extension for desktop computers (Chrome-family browsers first, Mac and Windows). You're on YouTube, a news article with an embedded video, or a social feed. You press a keyboard shortcut. The video pauses, the tool looks at the last several seconds along with the page context (the headline, the video title), and a description is spoken. You can then ask follow-up questions like "what does the sign say?" or "who just walked in?" without leaving the page. The same shortcut would work on images, or on a snapshot of the whole page. By default, descriptions would be spoken by your own screen reader, in your voice, at your speed, with your usual review commands, rather than by a separate voice the tool imposes.
How this differs from what exists today, as far as I can tell: tools like PiccyBot, Be My AI, and Seeing AI are excellent, but they need you to save or share the media into a separate app, and a lot of embedded web video can't easily be exported at all. Descriptions supplied by websites themselves only exist where the site owner chose to pay for them. I could not find anything that describes whatever you are already looking at, on demand, right where you are. If something like this already exists and I missed it, please tell me and I will go use that instead of building a duplicate.
My questions. Answer any of them, in any order:
1. Would you actually use this? Roughly how often do you run into web video with no description that you wished you could just ask about?
2. What do you use today when you hit an undescribed video, and where does it fall short?
3. Trust is my biggest worry. AI describers sometimes state wrong things with full confidence, and you have no way to check the video yourself. When the AI is not sure about something, which would you rather get: a guess clearly marked as a guess ("I think the sign says exit"), or nothing at all about that detail? Put another way: which bothers you more, a wrong detail or a missing one? And how do the tools you use today handle this?
4. Voice: would you rather descriptions come through your own screen reader, or through a separate natural-sounding voice? If your own screen reader: when a description is ready, should it wait until whatever is currently being read finishes, or interrupt right away?
5. Privacy: with cloud AI, snapshots of what you're watching leave your machine for processing. A fully local option is possible, but with lower quality on most hardware. How much does that tradeoff matter to you?
I have more specific questions saved up (how much visual detail you prefer, whole-page overviews for badly coded sites, remembering recurring characters in series you follow, how to handle live streams and video that cannot be paused, and whether Safari support matters to you), but I'll stop here. Whatever gets built will be free and open source, and I'll report back in this thread either way, including if the conclusion is "don't build it."
Thanks for reading.
Comments
My main use for video description.
I often see a lot of newsy videos that take ten minutes to listen to to get one main idea. I generally just grab the page and send it to gemini and ask for a summary. I don't know if the program you're considering would really make this sort of thing any easier in my use case or not.
I also use Gemini.
I just send a video to Gemini and ask it to give me a very detailed description of the video. If it's a very long YouTube video, I do the same here. I just attach the link and send it to Gemini and ask it to give me a detailed summary of the video. I don't see a big use case for this.
Thank you for feedback
I truly appreciate the feedback. This was one of things that I had considered, I just wasn't sure if there was any issues with this approach.
Cheers
JAWS Picture Smart
The one thing at least when I use Gemini, is it relies on the transcript. So you're not getting descriptions of visuals, such as a chart. JAWS's picture smart feature will take a picture of the screen and send it to AI. So if you can pause the video and keep the frame in view that works. I sometimes have had videos go blank when paused though, and then this doesn't work well. So if there was a more reliable way to say, hey I'm at this timestamp, grab the frames and AI them, being sure to actually get the frames. The other thing is if the tool could actually analyze the video, like a few seconds, instead of one still shot which is all JAWS can do. That may prove useful.
I do seem to recall some projects trying to do full video description via AI but can't think of the name. And in app form Seeing AI does have this feature though it is slow and does whole videos. (Disclaimer, it did, I've not checked in the past few months.) Hope this helps.
Video description
Gemini can output a full audio description text. What would be quite awesome, though, would be to get audio descriptions in realtime: the A.I. analyzes the video on the first pass to generate an audio description transcript using SMIL. Then it speaks that text synchronized with video playback. The extension might as well play the video itself. Wish I had this as a desktop app so I could use it on old DVDs. One day, probably sooner rather than alater, Copilot or something else will be able to just do it without an app or browser extension. Addendum: boy am I out of touch. SMIL has long since been supplanted by HTML 5 itself. Would love an accessible desktop media player to add audio description to a video, though--one that could accept a video URL as well. A simple accessible desktop app beats browser extension UX any day. At the risk of overgeneralizing, blind people are not fond of web apps.
Gemini
Maybe I'm the one out of date? How do you get Gemini to do audio description? Do you need a Pro subscription?
As for generalizing on web apps, that depends. I use them all the time. They can be designed well, and not well. Same as desktop apps. I do agree browsers make using many extensions difficult. I like the Gemini side panel in Chrome, but that is about it as far as extension type things go, and that panel isn't really an extension.
Re: DVDs
There are a handful of films that have yet to be audio described, professionally or otherwise.
I would kill to have said films described with the same level of accuracy and detail as some of the more accomplished (human) narrators.
I'm new to this ...
I'm new to this. Privacy is a main concern for me, and if the proposed add-on does what you, Jose, are describing and offer good factors of privacy (without involving a third party) I'd give it a try. Are you also considering making available descriptions in braille for those of us who are hard-of-hearing or prefer descriptions in writing rather than as audio output?
Thank you for asking what we want rather than assuming
I just wanted to pop on here to say thank you very much for coming on here and asking for our input as to what we would actually use rather than trying to build something and then just assuming that the people would want. I actually really don’t use anything with AI or description videos or anything like that for my computers, but I appreciate coming on here and actually asking for our input. There’s been a lot of assumptions made people would want hair blind and I think it’s a lot of times individual use cases. I just watched like regular YouTube videos or movies that I’ve already had body descriptions and that’s basically all I’ve really done. Movies with audio descriptions and that’s really kind of all I’ve done and then regular YouTube.
I would love something like that!
As previously stated, picture smart does have something like that but it’s not reliable. Besides, you have to basically pause every second so you can get a somewhat OK description. It gets tiring after two minutes.
How about different types of descriptions. Every few seconds, or per scene, or every minute? Or a summary if needed.
It would be super useful especially on YouTube where you can sometimes find old series or TV shows but there’s no audio description for those.
Follow up on new comments
Hi everyone,
First, sorry for the slow reply. I was away visiting family and only just got back. I honestly didn't expect this to go anywhere, so coming home to this many thoughtful responses was a real surprise. Thank you, all of you, this thread has already changed how I'm thinking about the whole idea.
Travis, your comment was the one that stopped me. You described the exact thing I was trying to get at, better than I did, that sending a video to an AI usually just gives you the words, and misses anything that's only on screen, like a chart. And that Picture Smart grabs a single frame that's often blank when the video's paused. That gap is the whole reason I was thinking about this, and hearing someone who clearly knows the space land on the same problem meant a lot.
Chris and Seamus, a real question, not a challenge, because you both know your own use better than I do. When you send a video to Gemini and it summarizes, has there ever been a time where something visual mattered, a chart, someone's reaction, text on the screen, and the summary quietly skipped it, and you only realized later? I'm trying to understand where a summary is genuinely enough, and where the picture is the thing you're missing.
Voracious, fair point on desktop apps versus browser extensions. I started with an extension because it seemed like the one spot nothing currently covers, but I hear you, and I'm not locked into that. Good to know.
Cordelia, yes. Text and braille output, not just speech, is something I'd want to include, and it looks like it's actually one of the easier things since screen readers already send text to a braille display. On privacy: my background is nearly three decades in IT and cybersecurity for corporations, so security and privacy aren't an afterthought for me, they're kind of baked into how I think. A local option that doesn't send anything to a third party is a core part of what I've been considering, not a bolt-on.
On the bigger idea a few of you raised, full, continuous description of whole movies, DVDs, old shows with no audio description, I'd love that too, honestly. Being straight with you: that's a much bigger project than one person can do well, so my thinking is to start small with the on-demand version and see if it's even useful before dreaming that big.
I'll also be honest about my own limits: I'm one person, not a developer by trade, so if this gets built it'll be in stages, the small, useful piece first, and the bigger pieces only if it earns its way there. I'd rather do one part well than promise the whole thing and disappear.
Last thing: the feedback here has been genuinely more useful than I expected, and it's shaped my thinking more than any amount of guessing on my own would have. If this problem sounds familiar to anyone you know in the community, I'd welcome their take too, the more real use cases I hear, the better I can figure out what's actually worth building.
I'm still just gathering input, not building anything yet, and I'll keep reporting back either way. Thank you again for taking a stranger's random idea seriously.
Jose (JDM)