AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (20.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Audio-guided implicit neural representation for local image stylization

Department of Artificial Intelligence, Korea University, Seoul 02841, Republic of Korea
NVIDIA Research, Nvidia Corporation, Santa Clara, CA 95051, USA
Graduate School of Culture Technology, KAIST, Seoul 34141, Republic of Korea
Hologram Research Center, Korea Electronics Technology Institute, Seoul 03924, Republic of Korea
Department of Computer Science and Engineering, Korea University, Seoul 02841, Republic of Korea

* Seung Hyun Lee and Sieun Kim contributed equally to this work.

Show Author Information

Abstract

We present a novel framework for audio-guided localized image stylization. Sound often provides information about the specific context of a scene and is closely related to a certain part of the scene or object. However, existing image stylization works have focused on stylizing the entire image using an image or text input. Stylizing a particular part of the image based on audio input is natural but challenging. This work proposes a framework in which a user provides an audio input to localize the target in the input image and another to locally stylize the target object or scene. We first produce a fine localization map using an audio-visual localization network leveraging CLIP embedding space. We then utilize an implicit neural representation (INR) along with the predicted localization map to stylize the target based on sound information. The INR manipulates local pixel values to be semantically consistent with the provided audio input. Our experiments show that the proposed framework outperforms other audio-guided stylization methods. Moreover, we observe that our method constructs concise localization maps and naturally manipulates the target object or scene in accordance with the given audio input.

Graphical Abstract

References

【1】
【1】
 
 
Computational Visual Media
Pages 1185-1204

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Lee SH, Kim S, Byeon W, et al. Audio-guided implicit neural representation for local image stylization. Computational Visual Media, 2024, 10(6): 1185-1204. https://doi.org/10.1007/s41095-024-0413-5

743

Views

51

Downloads

4

Crossref

4

Web of Science

5

Scopus

0

CSCD

Received: 10 August 2023
Accepted: 14 February 2024
Published: 14 August 2024
© The Author(s) 2024.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.

The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.

To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.

Other papers from this open access journal are available free of charge from http://www.springer.com/journal/41095. To submit a manuscript, please go to https://www.editorialmanager.com/cvmj.