Saved in:
| Main Authors: | Chen, Wei, Li, Yunan, Tian, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.06025 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
QRetinex-Net: Quaternion-Valued Retinex Decomposition for Low-Level Computer Vision Applications
by: Agaian, Sos, et al.
Published: (2025)
by: Agaian, Sos, et al.
Published: (2025)
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
WarmFed: Federated Learning with Warm-Start for Globalization and Personalization Via Personalized Diffusion Models
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
CEBSNet: Change-Excited and Background-Suppressed Network with Temporal Dependency Modeling for Bitemporal Change Detection
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
Flaws of ImageNet, Computer Vision's Favourite Dataset
by: Kisel, Nikita, et al.
Published: (2024)
by: Kisel, Nikita, et al.
Published: (2024)
GTSR: Subsurface Scattering Awared 3D Gaussians for Translucent Surface Reconstruction
by: Yuan, Youwen, et al.
Published: (2026)
by: Yuan, Youwen, et al.
Published: (2026)
Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
by: Zhu, Ye, et al.
Published: (2025)
by: Zhu, Ye, et al.
Published: (2025)
Lightweight Adapter Learning for More Generalized Remote Sensing Change Detection
by: Quan, Dou, et al.
Published: (2025)
by: Quan, Dou, et al.
Published: (2025)
ElasticLaneNet: An Efficient Geometry-Flexible Approach for Lane Detection
by: Feng, Yaxin, et al.
Published: (2023)
by: Feng, Yaxin, et al.
Published: (2023)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
by: Yuan, Zhenghang, et al.
Published: (2024)
by: Yuan, Zhenghang, et al.
Published: (2024)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
by: Zhang, Chengsheng, et al.
Published: (2026)
by: Zhang, Chengsheng, et al.
Published: (2026)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
MFDS-Net: Multi-Scale Feature Depth-Supervised Network for Remote Sensing Change Detection with Global Semantic and Detail Information
by: Huang, Zhenyang, et al.
Published: (2024)
by: Huang, Zhenyang, et al.
Published: (2024)
Deep Learning for Climate Action: Computer Vision Analysis of Visual Narratives on X
by: Prasse, Katharina, et al.
Published: (2025)
by: Prasse, Katharina, et al.
Published: (2025)
Neural Acquisition & Representation of Subsurface Scattering
by: Majumdar, Arjun, et al.
Published: (2026)
by: Majumdar, Arjun, et al.
Published: (2026)
Play to Generalize: Learning to Reason Through Game Play
by: Xie, Yunfei, et al.
Published: (2025)
by: Xie, Yunfei, et al.
Published: (2025)
Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
by: Peng, Shuai, et al.
Published: (2024)
by: Peng, Shuai, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Playing to Vision Foundation Model's Strengths in Stereo Matching
by: Liu, Chuang-Wei, et al.
Published: (2024)
by: Liu, Chuang-Wei, et al.
Published: (2024)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
On the Application of Egocentric Computer Vision to Industrial Scenarios
by: Chavan, Vivek, et al.
Published: (2024)
by: Chavan, Vivek, et al.
Published: (2024)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
by: Yakun, Cui, et al.
Published: (2026)
by: Yakun, Cui, et al.
Published: (2026)
Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
by: Yuan, Fan, et al.
Published: (2024)
by: Yuan, Fan, et al.
Published: (2024)
CCS: Clinical Consensus Selection for Radiology Report Generation
by: Zhang, Xi, et al.
Published: (2026)
by: Zhang, Xi, et al.
Published: (2026)
Caregiver Talk Shapes Toddler Vision: A Computational Study of Dyadic Play
by: Schaumlöffel, Timothy, et al.
Published: (2023)
by: Schaumlöffel, Timothy, et al.
Published: (2023)
Mitigating Biases in Surgical Operating Rooms with Geometry
by: Wang, Tony Danjun, et al.
Published: (2025)
by: Wang, Tony Danjun, et al.
Published: (2025)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024)
by: Xiong, Yuwen, et al.
Published: (2024)
MineNetCD: A Benchmark for Global Mining Change Detection on Remote Sensing Imagery
by: Yu, Weikang, et al.
Published: (2024)
by: Yu, Weikang, et al.
Published: (2024)
MultiTaskDeltaNet: Change Detection-based Image Segmentation for Operando ETEM with Application to Carbon Gasification Kinetics
by: Niu, Yushuo, et al.
Published: (2025)
by: Niu, Yushuo, et al.
Published: (2025)
Bridging Classical and Modern Computer Vision: PerceptiveNet for Tree Crown Semantic Segmentation
by: Voulgaris, Georgios
Published: (2025)
by: Voulgaris, Georgios
Published: (2025)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
by: He, Jingtao, et al.
Published: (2026)
by: He, Jingtao, et al.
Published: (2026)
Subsurface Scattering for 3D Gaussian Splatting
by: Dihlmann, Jan-Niklas, et al.
Published: (2024)
by: Dihlmann, Jan-Niklas, et al.
Published: (2024)
StrokeNet: Unveiling How to Learn Fine-Grained Interactions in Online Handwritten Stroke Classification
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
GS-Net: Generalizable Plug-and-Play 3D Gaussian Splatting Module
by: Zhang, Yichen, et al.
Published: (2024)
by: Zhang, Yichen, et al.
Published: (2024)
Similar Items
-
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
by: He, Xin, et al.
Published: (2025) -
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026) -
QRetinex-Net: Quaternion-Valued Retinex Decomposition for Low-Level Computer Vision Applications
by: Agaian, Sos, et al.
Published: (2025) -
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
by: Liu, Xuyang, et al.
Published: (2025) -
WarmFed: Federated Learning with Warm-Start for Globalization and Personalization Via Personalized Diffusion Models
by: Feng, Tao, et al.
Published: (2025)