Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xuesong, Li, Jia, Xu, Yunbo, Hu, Zhenzhen, Hong, Richang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
di: Xu, Yunbo, et al.
Pubblicazione: (2025)
di: Xu, Yunbo, et al.
Pubblicazione: (2025)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
di: Song, Zijie, et al.
Pubblicazione: (2023)
di: Song, Zijie, et al.
Pubblicazione: (2023)
Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
di: Wang, Xuesong, et al.
Pubblicazione: (2026)
di: Wang, Xuesong, et al.
Pubblicazione: (2026)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
di: Hou, Wenjin, et al.
Pubblicazione: (2026)
Seeing is not Believing: An Identity Hider for Human Vision Privacy Protection
di: Wang, Tao, et al.
Pubblicazione: (2023)
di: Wang, Tao, et al.
Pubblicazione: (2023)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2024)
di: Xiao, Jian, et al.
Pubblicazione: (2024)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
di: Cheng, Hao, et al.
Pubblicazione: (2025)
di: Cheng, Hao, et al.
Pubblicazione: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
di: He, Zhentao, et al.
Pubblicazione: (2025)
di: He, Zhentao, et al.
Pubblicazione: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
di: Song, Zijie, et al.
Pubblicazione: (2025)
di: Song, Zijie, et al.
Pubblicazione: (2025)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
di: Xuan, Weihao, et al.
Pubblicazione: (2025)
di: Xuan, Weihao, et al.
Pubblicazione: (2025)
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
di: Chen, Yin, et al.
Pubblicazione: (2024)
di: Chen, Yin, et al.
Pubblicazione: (2024)
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
di: Geng, Zibin, et al.
Pubblicazione: (2026)
di: Geng, Zibin, et al.
Pubblicazione: (2026)
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
di: Panchal, Utsav, et al.
Pubblicazione: (2025)
di: Panchal, Utsav, et al.
Pubblicazione: (2025)
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
di: Yu, Yangche, et al.
Pubblicazione: (2025)
di: Yu, Yangche, et al.
Pubblicazione: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
di: Liu, Ting, et al.
Pubblicazione: (2024)
di: Liu, Ting, et al.
Pubblicazione: (2024)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
di: Huang, Shaofei, et al.
Pubblicazione: (2025)
di: Huang, Shaofei, et al.
Pubblicazione: (2025)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
di: Liu, Zhining, et al.
Pubblicazione: (2025)
di: Liu, Zhining, et al.
Pubblicazione: (2025)
Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
di: Li, Jia, et al.
Pubblicazione: (2022)
di: Li, Jia, et al.
Pubblicazione: (2022)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
di: Ling, Chen, et al.
Pubblicazione: (2026)
di: Ling, Chen, et al.
Pubblicazione: (2026)
PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
di: Cui, Kai, et al.
Pubblicazione: (2025)
di: Cui, Kai, et al.
Pubblicazione: (2025)
DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
Believing is Seeing: Unobserved Object Detection using Generative Models
di: Bhattacharjee, Subhransu S., et al.
Pubblicazione: (2024)
di: Bhattacharjee, Subhransu S., et al.
Pubblicazione: (2024)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2025)
di: Xiao, Jian, et al.
Pubblicazione: (2025)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
di: Song, Zijie, et al.
Pubblicazione: (2023)
di: Song, Zijie, et al.
Pubblicazione: (2023)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
di: Wang, Xu, et al.
Pubblicazione: (2024)
di: Wang, Xu, et al.
Pubblicazione: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
di: Tan, Weiting, et al.
Pubblicazione: (2025)
di: Tan, Weiting, et al.
Pubblicazione: (2025)
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
di: Zhang, Huaxiang, et al.
Pubblicazione: (2024)
di: Zhang, Huaxiang, et al.
Pubblicazione: (2024)
SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding
di: Shi, Liangtao, et al.
Pubblicazione: (2025)
di: Shi, Liangtao, et al.
Pubblicazione: (2025)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
di: Liang, Dayong, et al.
Pubblicazione: (2025)
di: Liang, Dayong, et al.
Pubblicazione: (2025)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
di: Deng, Ailin, et al.
Pubblicazione: (2024)
di: Deng, Ailin, et al.
Pubblicazione: (2024)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
di: Fang, Zhenying, et al.
Pubblicazione: (2025)
di: Fang, Zhenying, et al.
Pubblicazione: (2025)
Image Captioning via Compact Bidirectional Architecture
di: Song, Zijie, et al.
Pubblicazione: (2022)
di: Song, Zijie, et al.
Pubblicazione: (2022)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
di: Hu, Xin, et al.
Pubblicazione: (2026)
di: Hu, Xin, et al.
Pubblicazione: (2026)
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
di: Zhou, Yuan, et al.
Pubblicazione: (2024)
di: Zhou, Yuan, et al.
Pubblicazione: (2024)
Enhancing Features in Long-tailed Data Using Large Vision Model
di: Han, Pengxiao, et al.
Pubblicazione: (2025)
di: Han, Pengxiao, et al.
Pubblicazione: (2025)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
di: Huang, Xin, et al.
Pubblicazione: (2025)
di: Huang, Xin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
di: Xu, Yunbo, et al.
Pubblicazione: (2025) -
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
di: Zhang, Xuesong, et al.
Pubblicazione: (2024) -
Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
di: Song, Zijie, et al.
Pubblicazione: (2023) -
Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
di: Wang, Xuesong, et al.
Pubblicazione: (2026) -
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
di: Hou, Wenjin, et al.
Pubblicazione: (2026)