GiVE: Guiding Visual Encoder to Perceive Overlooked Information
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Junjie, Ma, Jianghong, Zhang, Xiaofeng, Li, Yuhang, Shi, Jianyang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
por: He, Huiguo, et al.
Publicado: (2024)
por: He, Huiguo, et al.
Publicado: (2024)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
por: Song, Tianyi, et al.
Publicado: (2023)
por: Song, Tianyi, et al.
Publicado: (2023)
CountingFruit: Language-Guided 3D Fruit Counting with Semantic Gaussian Splatting
por: Li, Fengze, et al.
Publicado: (2025)
por: Li, Fengze, et al.
Publicado: (2025)
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
por: Wu, Linzhi, et al.
Publicado: (2024)
por: Wu, Linzhi, et al.
Publicado: (2024)
Decoupled Audio-Visual Dataset Distillation
por: Li, Wenyuan, et al.
Publicado: (2025)
por: Li, Wenyuan, et al.
Publicado: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
por: Wang, Yi, et al.
Publicado: (2025)
por: Wang, Yi, et al.
Publicado: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
por: Wang, Zhenyu, et al.
Publicado: (2026)
por: Wang, Zhenyu, et al.
Publicado: (2026)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
por: Liu, Che, et al.
Publicado: (2026)
por: Liu, Che, et al.
Publicado: (2026)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
por: Marinoni, Christian, et al.
Publicado: (2025)
por: Marinoni, Christian, et al.
Publicado: (2025)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
por: Li, Guangyao, et al.
Publicado: (2024)
por: Li, Guangyao, et al.
Publicado: (2024)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
por: Li, Ke, et al.
Publicado: (2026)
por: Li, Ke, et al.
Publicado: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
por: Cai, Dongnuan, et al.
Publicado: (2026)
por: Cai, Dongnuan, et al.
Publicado: (2026)
HaineiFRDM: Explore Diffusion to Restore Defects in Fast-Movement Films
por: Xun, Rongji, et al.
Publicado: (2025)
por: Xun, Rongji, et al.
Publicado: (2025)
Can I Trust Your Answer? Visually Grounded Video Question Answering
por: Xiao, Junbin, et al.
Publicado: (2023)
por: Xiao, Junbin, et al.
Publicado: (2023)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Audio Visual Segmentation Through Text Embeddings
por: Lee, Kyungbok, et al.
Publicado: (2025)
por: Lee, Kyungbok, et al.
Publicado: (2025)
Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLM
por: Li, Hongzhao, et al.
Publicado: (2024)
por: Li, Hongzhao, et al.
Publicado: (2024)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
por: Yi, Hongzhu, et al.
Publicado: (2026)
por: Yi, Hongzhu, et al.
Publicado: (2026)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
por: Zhang, Junzheng, et al.
Publicado: (2024)
por: Zhang, Junzheng, et al.
Publicado: (2024)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
por: Li, Fuhao, et al.
Publicado: (2026)
por: Li, Fuhao, et al.
Publicado: (2026)
One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation
por: Lu, Shuo, et al.
Publicado: (2026)
por: Lu, Shuo, et al.
Publicado: (2026)
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
por: Hong, Yuyang, et al.
Publicado: (2025)
por: Hong, Yuyang, et al.
Publicado: (2025)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
por: Luo, Bingjun, et al.
Publicado: (2025)
por: Luo, Bingjun, et al.
Publicado: (2025)
Attributes-aware Visual Emotion Representation Learning
por: Maharjan, Rahul Singh, et al.
Publicado: (2025)
por: Maharjan, Rahul Singh, et al.
Publicado: (2025)
Efficient Low-Resolution Face Recognition via Bridge Distillation
por: Ge, Shiming, et al.
Publicado: (2024)
por: Ge, Shiming, et al.
Publicado: (2024)
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
por: Zhang, Hanwen, et al.
Publicado: (2026)
por: Zhang, Hanwen, et al.
Publicado: (2026)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
por: Li, Yaowei, et al.
Publicado: (2025)
por: Li, Yaowei, et al.
Publicado: (2025)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
por: Xiao, Junbin, et al.
Publicado: (2025)
por: Xiao, Junbin, et al.
Publicado: (2025)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
por: Chen, Yi-Chun
Publicado: (2025)
por: Chen, Yi-Chun
Publicado: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
por: Zeng, Xiangyu, et al.
Publicado: (2024)
por: Zeng, Xiangyu, et al.
Publicado: (2024)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
por: Yu, Lijun, et al.
Publicado: (2023)
por: Yu, Lijun, et al.
Publicado: (2023)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
por: Zhou, Jinxing, et al.
Publicado: (2024)
por: Zhou, Jinxing, et al.
Publicado: (2024)
CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity
por: Manolache, Georgiana, et al.
Publicado: (2025)
por: Manolache, Georgiana, et al.
Publicado: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
por: Park, Kyu Ri, et al.
Publicado: (2024)
por: Park, Kyu Ri, et al.
Publicado: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
por: Ukai, Mahiro, et al.
Publicado: (2024)
por: Ukai, Mahiro, et al.
Publicado: (2024)
Ejemplares similares
-
BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
por: Zhou, Dongliang, et al.
Publicado: (2025) -
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
por: He, Huiguo, et al.
Publicado: (2024) -
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
por: Song, Tianyi, et al.
Publicado: (2023) -
CountingFruit: Language-Guided 3D Fruit Counting with Semantic Gaussian Splatting
por: Li, Fengze, et al.
Publicado: (2025) -
Landmark-Guided Cross-Speaker Lip Reading with Mutual Information Regularization
por: Wu, Linzhi, et al.
Publicado: (2024)