Global Context Compression with Interleaved Vision-Text Transformation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiao, Dian, Duan, Jiaxin, Zhao, Shuai, Leng, Jiabing, Zhang, Yiran, Huang, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
von: Wang, Bingli, et al.
Veröffentlicht: (2026)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
von: Che, Chang, et al.
Veröffentlicht: (2024)
von: Che, Chang, et al.
Veröffentlicht: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
von: Du, Hang, et al.
Veröffentlicht: (2025)
von: Du, Hang, et al.
Veröffentlicht: (2025)
Simple o3: Towards Interleaved Vision-Language Reasoning
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
GloTSFormer: Global Video Text Spotting Transformer
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
von: Zhang, Yizhen, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhen, et al.
Veröffentlicht: (2025)
Hierarchical Vision Transformer Enhanced by Graph Convolutional Network for Image Classification
von: Jiao, Haibin
Veröffentlicht: (2026)
von: Jiao, Haibin
Veröffentlicht: (2026)
Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
von: Xie, Roy, et al.
Veröffentlicht: (2026)
von: Xie, Roy, et al.
Veröffentlicht: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
von: Wu, Yike, et al.
Veröffentlicht: (2025)
von: Wu, Yike, et al.
Veröffentlicht: (2025)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
von: Zhang, Zhenguo, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenguo, et al.
Veröffentlicht: (2025)
Spiking Vision Transformer with Saccadic Attention
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2026)
von: Dai, Yang, et al.
Veröffentlicht: (2026)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
von: Li, Qingyun, et al.
Veröffentlicht: (2024)
von: Li, Qingyun, et al.
Veröffentlicht: (2024)
Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization
von: Snyder, Thomas, et al.
Veröffentlicht: (2026)
von: Snyder, Thomas, et al.
Veröffentlicht: (2026)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
von: Karmore, Aryan
Veröffentlicht: (2026)
von: Karmore, Aryan
Veröffentlicht: (2026)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
von: Li, Bing, et al.
Veröffentlicht: (2022)
von: Li, Bing, et al.
Veröffentlicht: (2022)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
von: Hu, Nanxing, et al.
Veröffentlicht: (2025)
von: Hu, Nanxing, et al.
Veröffentlicht: (2025)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers
von: Mei, Luoyu, et al.
Veröffentlicht: (2024)
von: Mei, Luoyu, et al.
Veröffentlicht: (2024)
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
Fine-Grained Cat Breed Recognition with Global Context Vision Transformer
von: Hera, Mowmita Parvin, et al.
Veröffentlicht: (2026)
von: Hera, Mowmita Parvin, et al.
Veröffentlicht: (2026)
LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
von: Jiang, Wen, et al.
Veröffentlicht: (2025)
von: Jiang, Wen, et al.
Veröffentlicht: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
Benchmarking Unlearning for Vision Transformers
von: Zhao, Kairan, et al.
Veröffentlicht: (2026)
von: Zhao, Kairan, et al.
Veröffentlicht: (2026)
A 2D Semantic-Aware Position Encoding for Vision Transformers
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
LAPX: Lightweight Hourglass Network with Global Context
von: Zhao, Haopeng, et al.
Veröffentlicht: (2025)
von: Zhao, Haopeng, et al.
Veröffentlicht: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024) -
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
von: Feng, Yukang, et al.
Veröffentlicht: (2025) -
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
von: Wang, Bingli, et al.
Veröffentlicht: (2026) -
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025) -
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)