FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Janghoon, Lee, Jungsoo, Hayat, Munawar, Hwang, Kyuwoong, Porikli, Fatih, Choi, Sungha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
von: Park, Hyojin, et al.
Veröffentlicht: (2026)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
von: Hedlin, Eric, et al.
Veröffentlicht: (2024)
Attention Guided Alignment in Efficient Vision-Language Models
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
von: Das, Debasmit, et al.
Veröffentlicht: (2025)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
von: Kadambi, Shreya, et al.
Veröffentlicht: (2025)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
von: Park, Sunghyun, et al.
Veröffentlicht: (2026)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
PosSAM: Panoptic Open-vocabulary Segment Anything
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
ToSA: Token Selective Attention for Efficient Vision Transformers
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Manish Kumar, et al.
Veröffentlicht: (2024)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
von: Li, Zhuojin, et al.
Veröffentlicht: (2026)
von: Li, Zhuojin, et al.
Veröffentlicht: (2026)
Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
Principles of Visual Tokens for Efficient Video Understanding
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
von: Bang, Jihwan, et al.
Veröffentlicht: (2025)
von: Bang, Jihwan, et al.
Veröffentlicht: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
DiffAugment: Diffusion based Long-Tailed Visual Relationship Recognition
von: Gupta, Parul, et al.
Veröffentlicht: (2024)
von: Gupta, Parul, et al.
Veröffentlicht: (2024)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
FouRA: Fourier Low Rank Adaptation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
von: Wang, Shida, et al.
Veröffentlicht: (2026)
von: Wang, Shida, et al.
Veröffentlicht: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
von: Lee, Inseo, et al.
Veröffentlicht: (2025)
von: Lee, Inseo, et al.
Veröffentlicht: (2025)
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
von: Shi, Yunxiao, et al.
Veröffentlicht: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
DuoLoRA : Cycle-consistent and Rank-disentangled Content-Style Personalization
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025) -
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025) -
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
von: Park, Sunghyun, et al.
Veröffentlicht: (2025) -
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
von: Park, Hyojin, et al.
Veröffentlicht: (2026) -
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
von: Park, Minho, et al.
Veröffentlicht: (2025)