SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Sicheng, Wang, Lintao, Zhu, Xiaogang, Lu, Xuequan, Wang, Zhiyong, Hu, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling
by: Li, Siran, et al.
Published: (2025)
by: Li, Siran, et al.
Published: (2025)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
by: Hu, Xuran, et al.
Published: (2026)
by: Hu, Xuran, et al.
Published: (2026)
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
by: Yu, Sicheng, et al.
Published: (2025)
by: Yu, Sicheng, et al.
Published: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026)
by: Han, Yudong, et al.
Published: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
by: Luo, Wang, et al.
Published: (2025)
by: Luo, Wang, et al.
Published: (2025)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
by: Liu, Jianming, et al.
Published: (2025)
by: Liu, Jianming, et al.
Published: (2025)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
A Recipe for Geometry-Aware 3D Mesh Transformers
by: Farazi, Mohammad, et al.
Published: (2024)
by: Farazi, Mohammad, et al.
Published: (2024)
YotoR-You Only Transform One Representation
by: Villa, José Ignacio Díaz, et al.
Published: (2024)
by: Villa, José Ignacio Díaz, et al.
Published: (2024)
Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing
by: Perera, Amal S., et al.
Published: (2026)
by: Perera, Amal S., et al.
Published: (2026)
Towards Hard and Soft Shadow Removal via Dual-Branch Separation Network and Vision Transformer
by: Liang, Jiajia
Published: (2025)
by: Liang, Jiajia
Published: (2025)
From CNNs to Transformers in Multimodal Human Action Recognition: A Survey
by: Shaikh, Muhammad Bilal, et al.
Published: (2024)
by: Shaikh, Muhammad Bilal, et al.
Published: (2024)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
by: Chen, Zhangquan, et al.
Published: (2026)
by: Chen, Zhangquan, et al.
Published: (2026)
FocusedAD: Character-centric Movie Audio Description
by: Ye, Xiaojun, et al.
Published: (2025)
by: Ye, Xiaojun, et al.
Published: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
by: Jian, Song, et al.
Published: (2025)
by: Jian, Song, et al.
Published: (2025)
A Challenging Benchmark of Anime Style Recognition
by: Li, Haotang, et al.
Published: (2022)
by: Li, Haotang, et al.
Published: (2022)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)
by: Han, Yudong, et al.
Published: (2024)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
by: Berjawi, Jad, et al.
Published: (2025)
by: Berjawi, Jad, et al.
Published: (2025)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
by: Zheng, Peng, et al.
Published: (2024)
by: Zheng, Peng, et al.
Published: (2024)
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
Interpretable Tau-PET Synthesis from Multimodal T1-Weighted and FLAIR MRI Using Partial Information Decomposition Guided Disentangled Quantized Half-UNet
by: Chopra, Agamdeep S., et al.
Published: (2026)
by: Chopra, Agamdeep S., et al.
Published: (2026)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
by: Hu, Chenggong, et al.
Published: (2026)
by: Hu, Chenggong, et al.
Published: (2026)
LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
by: Xiao, Yutong, et al.
Published: (2026)
by: Xiao, Yutong, et al.
Published: (2026)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026)
by: Tang, Guowei
Published: (2026)
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
by: Fang, Tiancheng, et al.
Published: (2026)
by: Fang, Tiancheng, et al.
Published: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
by: Xu, Zhenyu, et al.
Published: (2025)
by: Xu, Zhenyu, et al.
Published: (2025)
Topology-Aware Latent Diffusion for 3D Shape Generation
by: Hu, Jiangbei, et al.
Published: (2024)
by: Hu, Jiangbei, et al.
Published: (2024)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
by: Dubois, L'ea, et al.
Published: (2025)
by: Dubois, L'ea, et al.
Published: (2025)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
by: Zhao, Yiming
Published: (2026)
by: Zhao, Yiming
Published: (2026)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
FUSE-Flow: Scalable Real-Time Multi-View Point Cloud Reconstruction Using Confidence
by: Sun, Chentian
Published: (2026)
by: Sun, Chentian
Published: (2026)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
by: Tran, Duc Dang Trung, et al.
Published: (2024)
by: Tran, Duc Dang Trung, et al.
Published: (2024)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
by: He, Guangzhao, et al.
Published: (2025)
by: He, Guangzhao, et al.
Published: (2025)
GMAC: Global Multi-View Constraint for Automatic Multi-Camera Extrinsic Calibration
by: Sun, Chentian
Published: (2026)
by: Sun, Chentian
Published: (2026)
Similar Items
-
GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling
by: Li, Siran, et al.
Published: (2025) -
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
by: Hu, Xuran, et al.
Published: (2026) -
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
by: Yu, Sicheng, et al.
Published: (2025) -
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026) -
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)