Principles of Visual Tokens for Efficient Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hao, Xinyue, Li, Gen, Gowda, Shreyank N, Fisher, Robert B, Huang, Jonathan, Arnab, Anurag, Sevilla-Lara, Laura |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Telling Stories for Common Sense Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2023)
par: Gowda, Shreyank N, et autres
Publié: (2023)
Watt For What: Rethinking Deep Learning's Energy-Performance Relationship
par: Gowda, Shreyank N, et autres
Publié: (2023)
par: Gowda, Shreyank N, et autres
Publié: (2023)
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
par: Sathiyanarayanan, Shriram M, et autres
Publié: (2025)
par: Sathiyanarayanan, Shriram M, et autres
Publié: (2025)
Continual Learning Improves Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts
par: Kim, Kiyoon, et autres
Publié: (2024)
par: Kim, Kiyoon, et autres
Publié: (2024)
Reimagining Reality: A Comprehensive Survey of Video Inpainting Techniques
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
FE-Adapter: Adapting Image-based Emotion Classifiers to Videos
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
par: Gowda, Shreyank N, et autres
Publié: (2024)
par: Gowda, Shreyank N, et autres
Publié: (2024)
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
par: Gowda, Shreyank N, et autres
Publié: (2025)
par: Gowda, Shreyank N, et autres
Publié: (2025)
It's a Matter of Time: Three Lessons on Long-Term Motion for Perception
par: Davison, Willem, et autres
Publié: (2026)
par: Davison, Willem, et autres
Publié: (2026)
ZeroDiff++: Substantial Unseen Visual-semantic Correlation in Zero-shot Learning
par: Ye, Zihan, et autres
Publié: (2026)
par: Ye, Zihan, et autres
Publié: (2026)
Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks
par: Gahir, Amar, et autres
Publié: (2026)
par: Gahir, Amar, et autres
Publié: (2026)
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
par: Li, Gen, et autres
Publié: (2025)
par: Li, Gen, et autres
Publié: (2025)
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
par: Gowda, Shreyank N, et autres
Publié: (2025)
par: Gowda, Shreyank N, et autres
Publié: (2025)
ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
par: Ye, Zihan, et autres
Publié: (2024)
par: Ye, Zihan, et autres
Publié: (2024)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
par: Huang, Xiaohu, et autres
Publié: (2024)
par: Huang, Xiaohu, et autres
Publié: (2024)
Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
par: Gowda, Shreyank N, et autres
Publié: (2025)
par: Gowda, Shreyank N, et autres
Publié: (2025)
Compression as an Adversarial Amplifier Through Decision Space Reduction
par: Evans, Lewis, et autres
Publié: (2026)
par: Evans, Lewis, et autres
Publié: (2026)
Interpretable Zero-shot Learning with Infinite Class Concepts
par: Ye, Zihan, et autres
Publié: (2025)
par: Ye, Zihan, et autres
Publié: (2025)
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
par: Mustafa, Ahmed B, et autres
Publié: (2025)
par: Mustafa, Ahmed B, et autres
Publié: (2025)
Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
par: Mustafa, Ahmed B, et autres
Publié: (2026)
par: Mustafa, Ahmed B, et autres
Publié: (2026)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
par: Xiong, Tianwei, et autres
Publié: (2026)
par: Xiong, Tianwei, et autres
Publié: (2026)
Time-, Memory- and Parameter-Efficient Visual Adaptation
par: Mercea, Otniel-Bogdan, et autres
Publié: (2024)
par: Mercea, Otniel-Bogdan, et autres
Publié: (2024)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
par: Wu, Peiran, et autres
Publié: (2025)
par: Wu, Peiran, et autres
Publié: (2025)
Adversarial Robustness in Zero-Shot Learning:An Empirical Study on Class and Concept-Level Vulnerabilities
par: Peng, Zhiyuan, et autres
Publié: (2025)
par: Peng, Zhiyuan, et autres
Publié: (2025)
Dense Video Object Captioning from Disjoint Supervision
par: Zhou, Xingyi, et autres
Publié: (2023)
par: Zhou, Xingyi, et autres
Publié: (2023)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
par: Jiang, Jindong, et autres
Publié: (2025)
par: Jiang, Jindong, et autres
Publié: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
par: Kahatapitiya, Kumara, et autres
Publié: (2023)
par: Kahatapitiya, Kumara, et autres
Publié: (2023)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
par: Jin, Xinqi, et autres
Publié: (2025)
par: Jin, Xinqi, et autres
Publié: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
par: Zhang, Hongzhi, et autres
Publié: (2025)
par: Zhang, Hongzhi, et autres
Publié: (2025)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
par: Li, Gen, et autres
Publié: (2024)
par: Li, Gen, et autres
Publié: (2024)
Twin Trigger Generative Networks for Backdoor Attacks against Object Detection
par: Li, Zhiying, et autres
Publié: (2024)
par: Li, Zhiying, et autres
Publié: (2024)
SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
par: Liu, Hezhao, et autres
Publié: (2026)
par: Liu, Hezhao, et autres
Publié: (2026)
Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
par: Zhang, Chong, et autres
Publié: (2023)
par: Zhang, Chong, et autres
Publié: (2023)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
par: Zhang, Liang, et autres
Publié: (2024)
par: Zhang, Liang, et autres
Publié: (2024)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
par: Jain, Gagan, et autres
Publié: (2024)
par: Jain, Gagan, et autres
Publié: (2024)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
par: Zhang, Yunzhu, et autres
Publié: (2025)
par: Zhang, Yunzhu, et autres
Publié: (2025)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
par: Wang, Shida, et autres
Publié: (2026)
par: Wang, Shida, et autres
Publié: (2026)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
par: Cho, Janghoon, et autres
Publié: (2025)
par: Cho, Janghoon, et autres
Publié: (2025)
Documents similaires
-
Telling Stories for Common Sense Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2023) -
Watt For What: Rethinking Deep Learning's Energy-Performance Relationship
par: Gowda, Shreyank N, et autres
Publié: (2023) -
Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
par: Sathiyanarayanan, Shriram M, et autres
Publié: (2025) -
Continual Learning Improves Zero-Shot Action Recognition
par: Gowda, Shreyank N, et autres
Publié: (2024) -
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts
par: Kim, Kiyoon, et autres
Publié: (2024)