A blind woman with dark glasses holding a smartphone close to her face. Via gesellschaftsbilder.de ID 4212.
AI & Automation, Accessibility

Video for All: Scaling Accessibility with plain X

, 

Video has become the default language of modern media. However, for blind and visually disabled audiences, a clip without a proper description is hard to follow, if not useless: they hear the voices, but miss the action. At DW Innovation, we're trying to address the barrier with plain X, our very own AI-driven media production tool.

Have you come across the term AD? That's short for audio description. Wikipedia describes the concept as "a form of narration used to provide information surrounding key visual elements in a media work". To give you a couple of concrete examples: ADs tell you who has entered a room, what is written on a billboard, and how the scene changes. ADs make sure that people who can't see the picture can still follow the story. The catch: Writing good audio descriptions is intensive work that takes a lot of time and resources, which is why so few videos have them. 

Now at some point, subtitles went from a specialist extra to an everyday standard. We'd like ADs to make the same journey. Our long-term goal is that every single video can be accessed by everyone. A piece of software called plain X could be very helpful in this context.

plain X: a good starting point for AD

plain X is a media localization platform developed by DW in cooperation with Priberam. It brings together four workflows in one place: transcription, translation, subtitling and voice-over generation. plain x allows users to upload any video – and go from a raw to a localized version in no time.

This setup turns out to be a good fit for AD. Describing a scene, editing the text and turning it into speech is not so different from what editors already do with subtitles and voice-overs. So, rather than building a separate tool, we're adding an AD feature to a platform our editors already know and frequently use.

State-of-the-art AI for AD generation

Advances in AI-based image and scene recognition have made our accessibility goal very attainable. "AD for all, all the time!" is a slogan we can hopefully live up to in the near future.

Here's how the process works: plain X analyzes a video, detects the individual scenes, and inserts a suggested description for each one into the storyboard. An editor then reads through the suggestions and corrects or rewrites them. Based on that final version of the info text, the platform generates a synthetic voice-over.

Here's an example:

Screenshot of plain x interface
At a wooden table, the blindfolded curly-haired man in the red sweatshirt tastes candy – video processed by plain X. The platform detects the scene, picks what matters, writes the storyboard, and adds the voice-over. A human editor can refine anything.

At this point, it's important to be clear about one thing: We're not about replacing professional editors and web video producers. We want to use AI to do the heavy lifting. And we want our colleagues to keep control of the text, the description, the nuances. They have the final say on everything. 

We also see this as a testing ground, a place to find out what works, what doesn't, and where a machine misses what a human would notice.

Freezing the frame to make room for detail

At the moment, we mainly work with extended audio description (EAD). In this format, the video pauses, or "freezes", while the description is read out. This gives us time for descriptions that are extra useful, rather than squeezing a rushed voice-over into a two-second gap. Wherever possible, we fit descriptions into natural pauses in the dialogue, so the video keeps flowing. The pause is a fallback for when a complex scene needs more room.

We'd like to hear from you

Technology can remove barriers, but it can also create new ones, and those are easy to miss. Especially if you don't have to rely on specific tools yourself.

So, we're asking for feedback from the community, from people who use and benefit from ADs and EADs. Will our approach help? What could get in the way? What would you do differently? Don't hesitate to get in touch via innovation@dw.com.