DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Meng, Lingchen, Yang, Jianwei, Tian, Rui, Dai, Xiyang, Wu, Zuxuan, Gao, Jianfeng, Jiang, Yu-Gang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
por: Mo, Zhiwen, et al.
Publicado: (2026)
por: Mo, Zhiwen, et al.
Publicado: (2026)
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
por: Chen, Yitong, et al.
Publicado: (2024)
por: Chen, Yitong, et al.
Publicado: (2024)
Integrating Cybersecurity in Predictive Cost-Benefit Power Scheduling: A DeepStack Model with Dynamic Defense Mechanism
por: Peivand, Ali, et al.
Publicado: (2025)
por: Peivand, Ali, et al.
Publicado: (2025)
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
por: You, Zuyao, et al.
Publicado: (2025)
por: You, Zuyao, et al.
Publicado: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
por: Chen, Yitong, et al.
Publicado: (2025)
por: Chen, Yitong, et al.
Publicado: (2025)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
por: Wang, Junke, et al.
Publicado: (2024)
por: Wang, Junke, et al.
Publicado: (2024)
SEGIC: Unleashing the Emergent Correspondence for In-Context Segmentation
por: Meng, Lingchen, et al.
Publicado: (2023)
por: Meng, Lingchen, et al.
Publicado: (2023)
FOCUS: Towards Universal Foreground Segmentation
por: You, Zuyao, et al.
Publicado: (2025)
por: You, Zuyao, et al.
Publicado: (2025)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
por: Wang, Junke, et al.
Publicado: (2025)
por: Wang, Junke, et al.
Publicado: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
por: Zhou, Ziwei, et al.
Publicado: (2025)
por: Zhou, Ziwei, et al.
Publicado: (2025)
Effectiveness of Stacks in the Stacked Hilbert-Huang Transform
por: Lin, Lupin Chun-Che, et al.
Publicado: (2025)
por: Lin, Lupin Chun-Che, et al.
Publicado: (2025)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
por: Peng, Wujian, et al.
Publicado: (2024)
por: Peng, Wujian, et al.
Publicado: (2024)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
por: Chen, Yitong, et al.
Publicado: (2026)
por: Chen, Yitong, et al.
Publicado: (2026)
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
por: Wang, Junke, et al.
Publicado: (2023)
por: Wang, Junke, et al.
Publicado: (2023)
Stack-Sorting with Dotted-Pattern-Avoiding Stacks
por: Shieh, Hansen, et al.
Publicado: (2024)
por: Shieh, Hansen, et al.
Publicado: (2024)
Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
por: Gao, Shujian, et al.
Publicado: (2026)
por: Gao, Shujian, et al.
Publicado: (2026)
StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection
por: Hou, Yanning, et al.
Publicado: (2025)
por: Hou, Yanning, et al.
Publicado: (2025)
Stacked Triple Differences
por: Hsieh, Meng Hsuan
Publicado: (2026)
por: Hsieh, Meng Hsuan
Publicado: (2026)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
por: Mirzaei, Motahhare, et al.
Publicado: (2024)
Detection of HI filament: Pair Stacking vs. Filament Stacking
por: Meng, Yuxi, et al.
Publicado: (2025)
por: Meng, Yuxi, et al.
Publicado: (2025)
BASE: Burst-Adaptive Autoscaling via Stacked Ensembles for SLO Assurance and Cost Efficiency
por: Meng, Chunyang, et al.
Publicado: (2024)
por: Meng, Chunyang, et al.
Publicado: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
por: Xie, Yin, et al.
Publicado: (2024)
por: Xie, Yin, et al.
Publicado: (2024)
Learning Accurate Segmentation Purely from Self-Supervision
por: You, Zuyao, et al.
Publicado: (2026)
por: You, Zuyao, et al.
Publicado: (2026)
REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents
por: Tian, Rui, et al.
Publicado: (2024)
por: Tian, Rui, et al.
Publicado: (2024)
Stacking Fault Engineered Tensile‐Strained Delafossite CuAlO 2 for CO 2 Electroreduction to Deeply Reduced Products
por: Bin Liang, et al.
Publicado: (2026)
por: Bin Liang, et al.
Publicado: (2026)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
por: Yang, Sihan, et al.
Publicado: (2025)
por: Yang, Sihan, et al.
Publicado: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
por: Bytedance-Seed-Foundation-Code-Team, et al.
Publicado: (2024)
por: Bytedance-Seed-Foundation-Code-Team, et al.
Publicado: (2024)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
por: Hallinan, Skyler, et al.
Publicado: (2025)
por: Hallinan, Skyler, et al.
Publicado: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
por: Jain, Jitesh, et al.
Publicado: (2024)
por: Jain, Jitesh, et al.
Publicado: (2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
por: Xing, Zhen, et al.
Publicado: (2024)
por: Xing, Zhen, et al.
Publicado: (2024)
Full Stack Navigation, Mapping, and Planning for the Lunar Autonomy Challenge
por: Dai, Adam, et al.
Publicado: (2026)
por: Dai, Adam, et al.
Publicado: (2026)
ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation
por: Yang, Zihan, et al.
Publicado: (2026)
por: Yang, Zihan, et al.
Publicado: (2026)
Simple but Effective Compound Geometric Operations for Temporal Knowledge Graph Completion
por: Ying, Rui, et al.
Publicado: (2024)
por: Ying, Rui, et al.
Publicado: (2024)
Stack-sorting with Stacks Avoiding Vincular Patterns
por: Zhao, William
Publicado: (2024)
por: Zhao, William
Publicado: (2024)
Notes on Stack Machines and Quantum Stack Machines
por: Qiu, Daowen
Publicado: (2025)
por: Qiu, Daowen
Publicado: (2025)
Deep ReLU Networks Have Surprisingly Simple Polytopes
por: Fan, Feng-Lei, et al.
Publicado: (2023)
por: Fan, Feng-Lei, et al.
Publicado: (2023)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
por: Zhu, Xinyu, et al.
Publicado: (2025)
por: Zhu, Xinyu, et al.
Publicado: (2025)
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
por: Jiao, Pengkun, et al.
Publicado: (2025)
por: Jiao, Pengkun, et al.
Publicado: (2025)
ModelLock: Locking Your Model With a Spell
por: Gao, Yifeng, et al.
Publicado: (2024)
por: Gao, Yifeng, et al.
Publicado: (2024)
Deeply Supervised Skin Lesions Diagnosis with Stage and Branch Attention
por: Dai, Wei, et al.
Publicado: (2022)
por: Dai, Wei, et al.
Publicado: (2022)
Ejemplares similares
-
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
por: Mo, Zhiwen, et al.
Publicado: (2026) -
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
por: Chen, Yitong, et al.
Publicado: (2024) -
Integrating Cybersecurity in Predictive Cost-Benefit Power Scheduling: A DeepStack Model with Dynamic Defense Mechanism
por: Peivand, Ali, et al.
Publicado: (2025) -
Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
por: You, Zuyao, et al.
Publicado: (2025) -
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
por: Chen, Yitong, et al.
Publicado: (2025)