Similar Items
Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending
by: Ko, Junseok, et al.
Published: (2026)
by: Ko, Junseok, et al.
Published: (2026)
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
by: Lan, Yuqin, et al.
Published: (2026)
by: Lan, Yuqin, et al.
Published: (2026)
Determining Mosaic Resilience in Sugarcane Plants using Hyperspectral Images
by: Zia, Ali, et al.
Published: (2025)
by: Zia, Ali, et al.
Published: (2025)
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023)
by: Xie, Jiahao, et al.
Published: (2023)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
Rethinking Patient Education as Multi-turn Multi-modal Interaction
by: Yao, Zonghai, et al.
Published: (2026)
by: Yao, Zonghai, et al.
Published: (2026)
PhotoBot: Reference-Guided Interactive Photography via Natural Language
by: Limoyo, Oliver, et al.
Published: (2024)
by: Limoyo, Oliver, et al.
Published: (2024)
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence
by: Chen, Jiabin, et al.
Published: (2025)
by: Chen, Jiabin, et al.
Published: (2025)
Preserving Old Memories in Vivid Detail: Human-Interactive Photo Restoration Framework
by: Back, Seung-Yeon, et al.
Published: (2024)
by: Back, Seung-Yeon, et al.
Published: (2024)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
RITA: A Real-time Interactive Talking Avatars Framework
by: Cheng, Wuxinlin, et al.
Published: (2024)
by: Cheng, Wuxinlin, et al.
Published: (2024)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion
by: Li, Lijian
Published: (2025)
by: Li, Lijian
Published: (2025)
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
by: Liu, Shenxi, et al.
Published: (2025)
by: Liu, Shenxi, et al.
Published: (2025)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
by: Khalil, Ahmad, et al.
Published: (2025)
by: Khalil, Ahmad, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Interact3D: Compositional 3D Generation of Interactive Objects
by: Shan, Hui, et al.
Published: (2026)
by: Shan, Hui, et al.
Published: (2026)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
Robust Domain Generalization for Multi-modal Object Recognition
by: Qiao, Yuxin, et al.
Published: (2024)
by: Qiao, Yuxin, et al.
Published: (2024)
Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration
by: Xu, Jinglin, et al.
Published: (2026)
by: Xu, Jinglin, et al.
Published: (2026)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
A Review of Human-Object Interaction Detection
by: Wang, Yuxiao, et al.
Published: (2024)
by: Wang, Yuxiao, et al.
Published: (2024)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)
by: Zhang, Zeren, et al.
Published: (2024)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
by: He, Yuting, et al.
Published: (2026)
by: He, Yuting, et al.
Published: (2026)
Multi-modal Auto-regressive Modeling via Visual Words
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction
by: Pramov, Aleksandar
Published: (2025)
by: Pramov, Aleksandar
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
by: Chen, Hejia, et al.
Published: (2025)
by: Chen, Hejia, et al.
Published: (2025)
DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification
by: Zheng, Aihua, et al.
Published: (2026)
by: Zheng, Aihua, et al.
Published: (2026)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
by: Zhou, Zhaokun, et al.
Published: (2024)
by: Zhou, Zhaokun, et al.
Published: (2024)
Multi-modal Relation Distillation for Unified 3D Representation Learning
by: Wang, Huiqun, et al.
Published: (2024)
by: Wang, Huiqun, et al.
Published: (2024)
Culture-inspired Multi-modal Color Palette Generation and Colorization: A Chinese Youth Subculture Case
by: Li, Yufan, et al.
Published: (2021)
by: Li, Yufan, et al.
Published: (2021)
Towards Unified Multi-granularity Text Detection with Interactive Attention
by: Wan, Xingyu, et al.
Published: (2024)
by: Wan, Xingyu, et al.
Published: (2024)
ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
by: Min, Yue, et al.
Published: (2025)
by: Min, Yue, et al.
Published: (2025)
Similar Items
-
Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending
by: Ko, Junseok, et al.
Published: (2026) -
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
by: Lan, Yuqin, et al.
Published: (2026) -
Determining Mosaic Resilience in Sugarcane Plants using Hyperspectral Images
by: Zia, Ali, et al.
Published: (2025) -
MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation
by: Xie, Jiahao, et al.
Published: (2023) -
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
by: Wang, Haoming, et al.
Published: (2026)