Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yunbo, Zhang, Xuesong, Li, Jia, Hu, Zhenzhen, Hong, Richang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
Improving Out-of-Distribution Detection with Disentangled Foreground and Background Features
von: Ding, Choubo, et al.
Veröffentlicht: (2023)
von: Ding, Choubo, et al.
Veröffentlicht: (2023)
Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
DiSa: Saliency-Aware Foreground-Background Disentangled Framework for Open-Vocabulary Semantic Segmentation
von: Yao, Zhen, et al.
Veröffentlicht: (2026)
von: Yao, Zhen, et al.
Veröffentlicht: (2026)
Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
von: Xiao, Jian, et al.
Veröffentlicht: (2024)
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
von: Hu, Ming, et al.
Veröffentlicht: (2026)
von: Hu, Ming, et al.
Veröffentlicht: (2026)
Enhancing Few-Shot Out-of-Distribution Detection via the Refinement of Foreground and Background
von: Li, Tianyu, et al.
Veröffentlicht: (2026)
von: Li, Tianyu, et al.
Veröffentlicht: (2026)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
von: Liu, Jinlin, et al.
Veröffentlicht: (2024)
von: Liu, Jinlin, et al.
Veröffentlicht: (2024)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
von: Song, Zijie, et al.
Veröffentlicht: (2025)
von: Song, Zijie, et al.
Veröffentlicht: (2025)
Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
von: Yu, Yangche, et al.
Veröffentlicht: (2025)
von: Yu, Yangche, et al.
Veröffentlicht: (2025)
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
von: Chen, Yin, et al.
Veröffentlicht: (2024)
von: Chen, Yin, et al.
Veröffentlicht: (2024)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
von: Wu, Yuheng, et al.
Veröffentlicht: (2025)
Animal Identification with Independent Foreground and Background Modeling
von: Picek, Lukas, et al.
Veröffentlicht: (2024)
von: Picek, Lukas, et al.
Veröffentlicht: (2024)
Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
von: Tang, Shengeng, et al.
Veröffentlicht: (2024)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Multi-Scale Foreground-Background Confidence for Out-of-Distribution Segmentation
von: Marschall, Samuel, et al.
Veröffentlicht: (2024)
von: Marschall, Samuel, et al.
Veröffentlicht: (2024)
PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
von: Cui, Kai, et al.
Veröffentlicht: (2025)
von: Cui, Kai, et al.
Veröffentlicht: (2025)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
Demystifying Foreground-Background Memorization in Diffusion Models
von: Di, Jimmy Z., et al.
Veröffentlicht: (2025)
von: Di, Jimmy Z., et al.
Veröffentlicht: (2025)
Controllable Relation Disentanglement for Few-Shot Class-Incremental Learning
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
von: Xiao, Jian, et al.
Veröffentlicht: (2025)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents
von: Kang, Taewon
Veröffentlicht: (2026)
von: Kang, Taewon
Veröffentlicht: (2026)
Is Foreground Prototype Sufficient? Few-Shot Medical Image Segmentation with Background-Fused Prototype
von: Tang, Song, et al.
Veröffentlicht: (2024)
von: Tang, Song, et al.
Veröffentlicht: (2024)
What Happens Without Background? Constructing Foreground-Only Data for Fine-Grained Tasks
von: Wang, Yuetian, et al.
Veröffentlicht: (2024)
von: Wang, Yuetian, et al.
Veröffentlicht: (2024)
Boosting Latent Diffusion Models via Disentangled Representation Alignment
von: Page, John, et al.
Veröffentlicht: (2026)
von: Page, John, et al.
Veröffentlicht: (2026)
Localization-Guided Foreground Augmentation in Autonomous Driving
von: Yong, Jiawei, et al.
Veröffentlicht: (2026)
von: Yong, Jiawei, et al.
Veröffentlicht: (2026)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
von: Fang, Zhenying, et al.
Veröffentlicht: (2025)
von: Fang, Zhenying, et al.
Veröffentlicht: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
FoBa: A Foreground-Background co-Guided Method and New Benchmark for Remote Sensing Semantic Change Detection
von: Zhang, Haotian, et al.
Veröffentlicht: (2025)
von: Zhang, Haotian, et al.
Veröffentlicht: (2025)
SignAligner: Harmonizing Complementary Pose Modalities for Coherent Sign Language Generation
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation
von: Hu, Zixuan, et al.
Veröffentlicht: (2026)
von: Hu, Zixuan, et al.
Veröffentlicht: (2026)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
von: Chen, Yin, et al.
Veröffentlicht: (2023)
von: Chen, Yin, et al.
Veröffentlicht: (2023)
Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
von: Li, Jia, et al.
Veröffentlicht: (2022)
von: Li, Jia, et al.
Veröffentlicht: (2022)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
Iterative Adversarial Attack on Image-guided Story Ending Generation
von: Wang, Youze, et al.
Veröffentlicht: (2023)
von: Wang, Youze, et al.
Veröffentlicht: (2023)
Self-supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System
von: Dengxiong, Xiwen, et al.
Veröffentlicht: (2024)
von: Dengxiong, Xiwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024) -
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024) -
Improving Out-of-Distribution Detection with Disentangled Foreground and Background Features
von: Ding, Choubo, et al.
Veröffentlicht: (2023) -
Grid Jigsaw Representation with CLIP: A New Perspective on Image Clustering
von: Song, Zijie, et al.
Veröffentlicht: (2023) -
DiSa: Saliency-Aware Foreground-Background Disentangled Framework for Open-Vocabulary Semantic Segmentation
von: Yao, Zhen, et al.
Veröffentlicht: (2026)