Bottleneck Tokens for Unified Multimodal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Siyu, Ren, Jing, Liao, Zhaohe, Mao, Dongxiao, Ren, Xiangyuan, Zhang, Yiyi, Zhao, Haohua, Lin, Weixiong, Shaohua, Jiang, Zhang, Liqing, Zheng, Yuchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024)
by: Liao, Zhaohe, et al.
Published: (2024)
High-Quality 3D Head Reconstruction from Any Single Portrait Image
by: Zhang, Jianfu, et al.
Published: (2025)
by: Zhang, Jianfu, et al.
Published: (2025)
Memory Retrieval and Consolidation in Large Language Models through Function Tokens
by: Zhang, Shaohua, et al.
Published: (2025)
by: Zhang, Shaohua, et al.
Published: (2025)
LongRetriever: Towards Ultra-Long Sequence based Candidate Retrieval for Recommendation
by: Ren, Qin, et al.
Published: (2025)
by: Ren, Qin, et al.
Published: (2025)
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
by: Hu, Xinhao, et al.
Published: (2026)
by: Hu, Xinhao, et al.
Published: (2026)
Unified Multi-Domain Graph Pre-training for Homogeneous and Heterogeneous Graphs via Domain-Specific Expert Encoding
by: Liang, Chundong, et al.
Published: (2026)
by: Liang, Chundong, et al.
Published: (2026)
CHoE: Cross-Domain Heterogeneous Graph Prompt Learning via Structure-Conditioned Experts
by: Li, Peiyuan, et al.
Published: (2026)
by: Li, Peiyuan, et al.
Published: (2026)
MUG: Meta-path-aware Universal Heterogeneous Graph Pre-Training
by: Shan, Lianze, et al.
Published: (2026)
by: Shan, Lianze, et al.
Published: (2026)
LEDA: Latent Semantic Distribution Alignment for Multi-domain Graph Pre-training
by: Shan, Lianze, et al.
Published: (2026)
by: Shan, Lianze, et al.
Published: (2026)
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
by: Chen, Tianyi, et al.
Published: (2024)
by: Chen, Tianyi, et al.
Published: (2024)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
by: Xie, Jinheng, et al.
Published: (2024)
by: Xie, Jinheng, et al.
Published: (2024)
UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking
by: Ren, Qionglin, et al.
Published: (2025)
by: Ren, Qionglin, et al.
Published: (2025)
Graph Integrated Multimodal Concept Bottleneck Model
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Few-Shot Generalized Category Discovery With Retrieval-Guided Decision Boundary Enhancement
by: Ren, Yunhan, et al.
Published: (2025)
by: Ren, Yunhan, et al.
Published: (2025)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
by: Wu, Qiong, et al.
Published: (2024)
by: Wu, Qiong, et al.
Published: (2024)
Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models
by: Li, Xiaohe, et al.
Published: (2026)
by: Li, Xiaohe, et al.
Published: (2026)
dnaGrinder: a lightweight and high-capacity genomic foundation model
by: Zhao, Qihang, et al.
Published: (2024)
by: Zhao, Qihang, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
by: Zhang, Peiying, et al.
Published: (2025)
by: Zhang, Peiying, et al.
Published: (2025)
CLASS: Enhancing Cross-Modal Text-Molecule Retrieval Performance and Training Efficiency
by: Wu, Hongyan, et al.
Published: (2025)
by: Wu, Hongyan, et al.
Published: (2025)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
From Cancer Drivers to Cancer Keepers: Paradigm Shift and Clinical Implications
by: Zhang, Xizhe, et al.
Published: (2025)
by: Zhang, Xizhe, et al.
Published: (2025)
Refined HLA Linkage Disequilibrium Architectures of World Populations by a Novel Allelic Correlation Measure
by: Zhang, Fei, et al.
Published: (2025)
by: Zhang, Fei, et al.
Published: (2025)
MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
by: Xiang, Mochu, et al.
Published: (2026)
by: Xiang, Mochu, et al.
Published: (2026)
Editable Concept Bottleneck Models
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
by: Jiang, Yutao, et al.
Published: (2025)
by: Jiang, Yutao, et al.
Published: (2025)
Alternating and Symmetric Separability of Free Products
by: Zhao, Dongxiao, et al.
Published: (2026)
by: Zhao, Dongxiao, et al.
Published: (2026)
Howson groups which are not strongly Howson
by: Zhang, Qiang, et al.
Published: (2024)
by: Zhang, Qiang, et al.
Published: (2024)
A note on test elements for monomorphisms of free groups
by: Zhao, Dongxiao, et al.
Published: (2024)
by: Zhao, Dongxiao, et al.
Published: (2024)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Exploring Dynamic Properties of Backdoor Training Through Information Bottleneck
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
by: Lu, Yanzuo, et al.
Published: (2025)
by: Lu, Yanzuo, et al.
Published: (2025)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
by: Li, Wenyun, et al.
Published: (2025)
by: Li, Wenyun, et al.
Published: (2025)
Transferable Adversarial Face Attack with Text Controlled Attribute
by: Li, Wenyun, et al.
Published: (2024)
by: Li, Wenyun, et al.
Published: (2024)
Token Bottleneck: One Token to Remember Dynamics
by: Kim, Taekyung, et al.
Published: (2025)
by: Kim, Taekyung, et al.
Published: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Similar Items
-
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024) -
High-Quality 3D Head Reconstruction from Any Single Portrait Image
by: Zhang, Jianfu, et al.
Published: (2025) -
Memory Retrieval and Consolidation in Large Language Models through Function Tokens
by: Zhang, Shaohua, et al.
Published: (2025) -
LongRetriever: Towards Ultra-Long Sequence based Candidate Retrieval for Recommendation
by: Ren, Qin, et al.
Published: (2025) -
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)