TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Tingyu, Li, Mingxiao, Tuytelaars, Tinne, Moens, Marie-Francine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023)
by: Qu, Tingyu, et al.
Published: (2023)
Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks
by: Qu, Tingyu, et al.
Published: (2024)
by: Qu, Tingyu, et al.
Published: (2024)
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024)
by: Li, Mingxiao, et al.
Published: (2024)
DM-Align: Leveraging the Power of Natural Language Instructions to Make Changes to Images
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Alleviating Exposure Bias in Diffusion Models through Sampling with Shifted Time Steps
by: Li, Mingxiao, et al.
Published: (2023)
by: Li, Mingxiao, et al.
Published: (2023)
OASIS: Online Sample Selection for Continual Visual Instruction Tuning
by: Lee, Minjae, et al.
Published: (2025)
by: Lee, Minjae, et al.
Published: (2025)
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
by: Lou, Haoran, et al.
Published: (2025)
by: Lou, Haoran, et al.
Published: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
by: Shen, Leqi, et al.
Published: (2025)
by: Shen, Leqi, et al.
Published: (2025)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
by: Lu, Weiheng, et al.
Published: (2024)
by: Lu, Weiheng, et al.
Published: (2024)
NeuroCine: Decoding Vivid Video Sequences from Human Brain Activties
by: Sun, Jingyuan, et al.
Published: (2024)
by: Sun, Jingyuan, et al.
Published: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024)
by: Zeer, Ahmed, et al.
Published: (2024)
Remembering by Reconstructing: Domain Incremental Learning With Test-Time Training on Video Streams
by: Swinnen, Jonathan, et al.
Published: (2026)
by: Swinnen, Jonathan, et al.
Published: (2026)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
by: Liang, Han, et al.
Published: (2024)
by: Liang, Han, et al.
Published: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
by: Lin, Bin, et al.
Published: (2023)
by: Lin, Bin, et al.
Published: (2023)
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
by: Yan, Dawei, et al.
Published: (2024)
by: Yan, Dawei, et al.
Published: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
Action-based image editing guided by human instructions
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
PEO: Training-Free Aesthetic Quality Enhancement in Pre-Trained Text-to-Image Diffusion Models with Prompt Embedding Optimization
by: Margaryan, Hovhannes, et al.
Published: (2025)
by: Margaryan, Hovhannes, et al.
Published: (2025)
Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
by: Zamini, Mohamad, et al.
Published: (2025)
by: Zamini, Mohamad, et al.
Published: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
by: Inal, Gokce, et al.
Published: (2026)
by: Inal, Gokce, et al.
Published: (2026)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
by: Araujo, Vladimir, et al.
Published: (2024)
by: Araujo, Vladimir, et al.
Published: (2024)
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound
by: Guo, Xuechen, et al.
Published: (2024)
by: Guo, Xuechen, et al.
Published: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
Similar Items
-
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
by: Li, Mingxiao, et al.
Published: (2025) -
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023) -
Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks
by: Qu, Tingyu, et al.
Published: (2024) -
Animate Your Motion: Turning Still Images into Dynamic Videos
by: Li, Mingxiao, et al.
Published: (2024) -
DM-Align: Leveraging the Power of Natural Language Instructions to Make Changes to Images
by: Trusca, Maria Mihaela, et al.
Published: (2024)