Abstract
Additionally the continued rise of generative artificial intelligence (GenAI) is leading to the easy creation of believable text, images, audio, and 3D virtual assets with which to populate the metaverse. However, the growing blurring of indistinctness between human-made and machine-made content has caused serious questions about legitimate content, intellectual property (IP) copyright and ethical transparency. This paper proposes a Multimodal AI-Watermarking Framework (MAI-WF) to protect generative content in the visual, auditory, and textual modalities in immersive metaverse interactions. The proposed framework includes deep neural embeddings, watermarking for reversibility of watermarked content and cross-modal feature fusion in order to provide authenticity verification & ownership tracking without a perceptual quality degradation. Experiments on publicly available datasets (MS-COCO, VCTK, and WikiText-103) show that the average Peak Signal-to-Noise Ratio (PSNR) for images, Signal-to-Noise Ratio (SNR) for audio, and Bit-Error Rate (BER) less than 1% are obtained in text embedding by MAI-WF with a robustness against compression, scaling and adversarial transformations. The framework supports interoperability with identity registries that are based on blockchains for providing traceability in decentralised metaverse environments.
Author supplied keywords
Cite
CITATION STYLE
Dixit, A., Gupta, A. K., Saxena, S., Midhunchakkaravarthy, D., & Gupta, D. (2026). Multimodal AI-Watermarking for Protecting Generative Content in Metaverse Interactions. Metaverse, 7(2). https://doi.org/10.54517/m8436
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.