Saved in:
| Main Authors: | Xia, Weihao, de Charette, Raoul, Öztireli, Cengiz, Xue, Jing-Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.07202 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multigranular Evaluation for Brain Visual Decoding
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
DREAM: Visual Decoding from Reversing Human Visual System
by: Xia, Weihao, et al.
Published: (2023)
by: Xia, Weihao, et al.
Published: (2023)
Exploring The Visual Feature Space for Multimodal Neural Decoding
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
RETRO: REthinking Tactile Representation Learning with Material PriOrs
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
by: Cao, Anh-Quan, et al.
Published: (2023)
by: Cao, Anh-Quan, et al.
Published: (2023)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
by: Zhou, Chenliang, et al.
Published: (2022)
by: Zhou, Chenliang, et al.
Published: (2022)
DenseMTL: Cross-task Attention Mechanism for Dense Multi-task Learning
by: Lopes, Ivan, et al.
Published: (2022)
by: Lopes, Ivan, et al.
Published: (2022)
StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
by: Cao, Anh-Quan, et al.
Published: (2025)
by: Cao, Anh-Quan, et al.
Published: (2025)
A Survey on Text-Driven 360-Degree Panorama Generation
by: Wang, Hai, et al.
Published: (2025)
by: Wang, Hai, et al.
Published: (2025)
Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction
by: Yang, Xiaoli, et al.
Published: (2026)
by: Yang, Xiaoli, et al.
Published: (2026)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025)
by: Lu, Xingyuan, et al.
Published: (2025)
BrainChat: Decoding Semantic Information from fMRI using Vision-language Pretrained Models
by: Huang, Wanaiu
Published: (2024)
by: Huang, Wanaiu
Published: (2024)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
by: Chen, Xiaokang, et al.
Published: (2025)
by: Chen, Xiaokang, et al.
Published: (2025)
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
by: Zheng, Sipeng, et al.
Published: (2024)
by: Zheng, Sipeng, et al.
Published: (2024)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
by: Ma, Yiyang, et al.
Published: (2024)
by: Ma, Yiyang, et al.
Published: (2024)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
by: Lee, Segyu, et al.
Published: (2026)
by: Lee, Segyu, et al.
Published: (2026)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
by: Jiang, Lingjie, et al.
Published: (2025)
by: Jiang, Lingjie, et al.
Published: (2025)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
by: Ren, Lingfeng, et al.
Published: (2026)
by: Ren, Lingfeng, et al.
Published: (2026)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
by: Liang, Jiafeng, et al.
Published: (2025)
by: Liang, Jiafeng, et al.
Published: (2025)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
by: Hong, Jixiang, et al.
Published: (2025)
by: Hong, Jixiang, et al.
Published: (2025)
Quartet of Diffusions: Structure-Aware Point Cloud Generation through Part and Symmetry Guidance
by: Zhou, Chenliang, et al.
Published: (2026)
by: Zhou, Chenliang, et al.
Published: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
by: Ji, Yicheng, et al.
Published: (2025)
by: Ji, Yicheng, et al.
Published: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
by: Mañas, Oscar, et al.
Published: (2025)
by: Mañas, Oscar, et al.
Published: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
by: Yu, Weihao, et al.
Published: (2023)
by: Yu, Weihao, et al.
Published: (2023)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
by: Xie, Xudong, et al.
Published: (2024)
by: Xie, Xudong, et al.
Published: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Similar Items
-
Multigranular Evaluation for Brain Visual Decoding
by: Xia, Weihao, et al.
Published: (2025) -
DREAM: Visual Decoding from Reversing Human Visual System
by: Xia, Weihao, et al.
Published: (2023) -
Exploring The Visual Feature Space for Multimodal Neural Decoding
by: Xia, Weihao, et al.
Published: (2025) -
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024) -
RETRO: REthinking Tactile Representation Learning with Material PriOrs
by: Xia, Weihao, et al.
Published: (2025)