ReconBoost: Boosting Can Achieve Modality Reconcilement
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Cong, Xu, Qianqian, Bao, Shilong, Yang, Zhiyong, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text Encoders
by: Han, Boyu, et al.
Published: (2025)
by: Han, Boyu, et al.
Published: (2025)
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2025)
by: Han, Boyu, et al.
Published: (2025)
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
by: Shang, Ziqiao, et al.
Published: (2023)
by: Shang, Ziqiao, et al.
Published: (2023)
Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
by: Bao, Shilong, et al.
Published: (2025)
by: Bao, Shilong, et al.
Published: (2025)
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
AUCSeg: AUC-oriented Pixel-level Long-tail Semantic Segmentation
by: Han, Boyu, et al.
Published: (2024)
by: Han, Boyu, et al.
Published: (2024)
Regularized Contrastive Partial Multi-view Outlier Detection
by: Wang, Yijia, et al.
Published: (2024)
by: Wang, Yijia, et al.
Published: (2024)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
by: Liu, Che, et al.
Published: (2026)
by: Liu, Che, et al.
Published: (2026)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
by: Lin, Xiang, et al.
Published: (2025)
by: Lin, Xiang, et al.
Published: (2025)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
by: Qin, Jie, et al.
Published: (2025)
by: Qin, Jie, et al.
Published: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
by: Yu, Lijun
Published: (2024)
by: Yu, Lijun
Published: (2024)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification on the DAIC-WOZ
by: Patapati, Santosh V.
Published: (2024)
by: Patapati, Santosh V.
Published: (2024)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
Flow Generator Matching
by: Huang, Zemin, et al.
Published: (2024)
by: Huang, Zemin, et al.
Published: (2024)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024)
by: Li, Guangyao, et al.
Published: (2024)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
by: Zhou, Shengli, et al.
Published: (2026)
by: Zhou, Shengli, et al.
Published: (2026)
HOIN: High-Order Implicit Neural Representations
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation
by: Pan, Yining, et al.
Published: (2025)
by: Pan, Yining, et al.
Published: (2025)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
by: Huang, Po-Hsuan, et al.
Published: (2024)
by: Huang, Po-Hsuan, et al.
Published: (2024)
Deep ReLU Networks Have Surprisingly Simple Polytopes
by: Fan, Feng-Lei, et al.
Published: (2023)
by: Fan, Feng-Lei, et al.
Published: (2023)
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
by: Pan, Li, et al.
Published: (2025)
by: Pan, Li, et al.
Published: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations
by: Yang, Zhiyong, et al.
Published: (2024)
by: Yang, Zhiyong, et al.
Published: (2024)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
OneLLM: One Framework to Align All Modalities with Language
by: Han, Jiaming, et al.
Published: (2023)
by: Han, Jiaming, et al.
Published: (2023)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
by: Yang, Chenglin, et al.
Published: (2023)
by: Yang, Chenglin, et al.
Published: (2023)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
by: Zhang, Shiyi, et al.
Published: (2026)
by: Zhang, Shiyi, et al.
Published: (2026)
Bidirectional Logits Tree: Pursuing Granularity Reconcilement in Fine-Grained Classification
by: Lu, Zhiguang, et al.
Published: (2024)
by: Lu, Zhiguang, et al.
Published: (2024)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Can We Edit Multimodal Large Language Models?
by: Cheng, Siyuan, et al.
Published: (2023)
by: Cheng, Siyuan, et al.
Published: (2023)
HLL: Can Agents Cross Humanity's Last Line of Verification?
by: Song, Xinhao, et al.
Published: (2026)
by: Song, Xinhao, et al.
Published: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Similar Items
-
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024) -
LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text Encoders
by: Han, Boyu, et al.
Published: (2025) -
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2025) -
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2026) -
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
by: Shang, Ziqiao, et al.
Published: (2023)