Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Kai, Guo, Jianyuan, Tang, Yehui, He, Wei, Wu, Enhua, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ParameterNet: Parameters Are All You Need
von: Han, Kai, et al.
Veröffentlicht: (2023)
von: Han, Kai, et al.
Veröffentlicht: (2023)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
von: Song, Baiyang, et al.
Veröffentlicht: (2026)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
von: Wang, Chengcheng, et al.
Veröffentlicht: (2024)
von: Wang, Chengcheng, et al.
Veröffentlicht: (2024)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
von: Hao, Zhiwei, et al.
Veröffentlicht: (2025)
von: Hao, Zhiwei, et al.
Veröffentlicht: (2025)
GhostNetV3: Exploring the Training Strategies for Compact Models
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
von: Ding, Ning, et al.
Veröffentlicht: (2023)
von: Ding, Ning, et al.
Veröffentlicht: (2023)
A Survey on Transformer Compression
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
Gold-YOLO: Efficient Object Detector via Gather-and-Distribute Mechanism
von: Wang, Chengcheng, et al.
Veröffentlicht: (2023)
von: Wang, Chengcheng, et al.
Veröffentlicht: (2023)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
Data-efficient Large Vision Models through Sequential Autoregression
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
Fine-gained Zero-shot Video Sampling
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
von: Chen, Dengsheng, et al.
Veröffentlicht: (2024)
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
RESP: Reference-guided Sequential Prompting for Visual Glitch Detection in Video Games
von: Yu, Yakun, et al.
Veröffentlicht: (2026)
von: Yu, Yakun, et al.
Veröffentlicht: (2026)
FreeViS: Training-free Video Stylization with Inconsistent References
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
Space-time Reinforcement Network for Video Object Segmentation
von: Chen, Yadang, et al.
Veröffentlicht: (2024)
von: Chen, Yadang, et al.
Veröffentlicht: (2024)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
von: Guo, Pinxue, et al.
Veröffentlicht: (2024)
LightCtrl: Training-free Controllable Video Relighting
von: Peng, Yizuo, et al.
Veröffentlicht: (2026)
von: Peng, Yizuo, et al.
Veröffentlicht: (2026)
Training-Free Robust Interactive Video Object Segmentation
von: Wei, Xiaoli, et al.
Veröffentlicht: (2024)
von: Wei, Xiaoli, et al.
Veröffentlicht: (2024)
PhyRPR: Training-Free Physics-Constrained Video Generation
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
von: Shu, Han, et al.
Veröffentlicht: (2023)
von: Shu, Han, et al.
Veröffentlicht: (2023)
How Important are Videos for Training Video LLMs?
von: Lydakis, George, et al.
Veröffentlicht: (2025)
von: Lydakis, George, et al.
Veröffentlicht: (2025)
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
von: Shen, Leqi, et al.
Veröffentlicht: (2025)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
von: Xie, Chaohao, et al.
Veröffentlicht: (2025)
von: Xie, Chaohao, et al.
Veröffentlicht: (2025)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos
von: Sun, Jiakai, et al.
Veröffentlicht: (2024)
von: Sun, Jiakai, et al.
Veröffentlicht: (2024)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
von: Wu, Zhenghuang, et al.
Veröffentlicht: (2026)
von: Wu, Zhenghuang, et al.
Veröffentlicht: (2026)
FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2026)
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ParameterNet: Parameters Are All You Need
von: Han, Kai, et al.
Veröffentlicht: (2023) -
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
von: Song, Baiyang, et al.
Veröffentlicht: (2026) -
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024) -
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024) -
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
von: Wang, Chengcheng, et al.
Veröffentlicht: (2024)