CoS: Chain-of-Shot Prompting for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Jian, Cheng, Zixu, Si, Chenyang, Li, Wei, Gong, Shaogang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
von: Hu, Jian, et al.
Veröffentlicht: (2025)
von: Hu, Jian, et al.
Veröffentlicht: (2025)
Grounding Video Reasoning in Physical Signals
von: Osmanli, Alibay, et al.
Veröffentlicht: (2026)
von: Osmanli, Alibay, et al.
Veröffentlicht: (2026)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
von: Hu, Jian, et al.
Veröffentlicht: (2025)
von: Hu, Jian, et al.
Veröffentlicht: (2025)
Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable Segmentation
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
InvSeg: Test-Time Prompt Inversion for Semantic Segmentation
von: Lin, Jiayi, et al.
Veröffentlicht: (2024)
von: Lin, Jiayi, et al.
Veröffentlicht: (2024)
Few-Shot Image Generation by Conditional Relaxing Diffusion Inversion
von: Cao, Yu, et al.
Veröffentlicht: (2024)
von: Cao, Yu, et al.
Veröffentlicht: (2024)
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
von: Hu, Jian, et al.
Veröffentlicht: (2025)
von: Hu, Jian, et al.
Veröffentlicht: (2025)
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding
von: Cheng, Zixu, et al.
Veröffentlicht: (2024)
von: Cheng, Zixu, et al.
Veröffentlicht: (2024)
Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
von: Cao, Yu, et al.
Veröffentlicht: (2025)
von: Cao, Yu, et al.
Veröffentlicht: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
Zero-Shot Long-Form Video Understanding through Screenplay
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
von: Cai, Weitong, et al.
Veröffentlicht: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2025)
von: Jahagirdar, Soumya Shamarao, et al.
Veröffentlicht: (2025)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Spatial Reasoning in Segmentation
von: Lin, Jiayi, et al.
Veröffentlicht: (2025)
von: Lin, Jiayi, et al.
Veröffentlicht: (2025)
Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking
von: Sun, Shitong, et al.
Veröffentlicht: (2023)
von: Sun, Shitong, et al.
Veröffentlicht: (2023)
CoDA: Instructive Chain-of-Domain Adaptation with Severity-Aware Visual Prompt Tuning
von: Gong, Ziyang, et al.
Veröffentlicht: (2024)
von: Gong, Ziyang, et al.
Veröffentlicht: (2024)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
von: Yang, Ying, et al.
Veröffentlicht: (2026)
von: Yang, Ying, et al.
Veröffentlicht: (2026)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
von: Sun, Shangkun, et al.
Veröffentlicht: (2024)
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection
von: Zeng, Fanhu, et al.
Veröffentlicht: (2024)
von: Zeng, Fanhu, et al.
Veröffentlicht: (2024)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection
von: Chen, Qiyu, et al.
Veröffentlicht: (2025)
von: Chen, Qiyu, et al.
Veröffentlicht: (2025)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
von: Wu, Zhixuan, et al.
Veröffentlicht: (2026)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
von: Mo, Ye, et al.
Veröffentlicht: (2025)
von: Mo, Ye, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025) -
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
von: Hu, Jian, et al.
Veröffentlicht: (2025) -
Grounding Video Reasoning in Physical Signals
von: Osmanli, Alibay, et al.
Veröffentlicht: (2026) -
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026) -
ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
von: Hu, Jian, et al.
Veröffentlicht: (2025)