Artificial Intelligence in Multimedia Content Generation: A Review of Audio and Video Synthesis Techniques

5Citations
Citations of this article
21Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Recent breakthroughs in generative AI have markedly elevated the realism and controllability of synthetic media. In the visual modality, long-context attention mechanisms and diffusion-style refinements now deliver videos with superior temporal consistency, spatial coherence, and high-resolution detail. These techniques underpin an expanding set of applications ranging from text-guided storyboarding and animation to engineering visualization and virtual prototyping. In the audio modality, token-based representations combined with hierarchical decoding enable the direct production of faithful speech, music, and ambient sound from textual prompts, powering rapid voice-over creation, personalized music, and immersive soundscapes. The frontier is shifting toward unified audio–visual pipelines that synchronize imagery with dialog, sound effects, and ambience, promising end-to-end tooling for a wide variety of applications such as education, simulation, entertainment, and accessible content production. This review surveys these advances across modalities and outlines future research directions focused on improving generation efficiency, coherence, and controllability across modalities.

Cite

CITATION STYLE

APA

Ding, C., & Bhowmik, R. (2026). Artificial Intelligence in Multimedia Content Generation: A Review of Audio and Video Synthesis Techniques. Journal of the Society for Information Display, 34(2), 49–67. https://doi.org/10.1002/jsid.2111

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free