Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Haoyu, Liu, Meng, Feng, Yisen, Wang, Yaowei, Guan, Weili, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
OSGNet @ Ego4D Episodic Memory Challenge 2025
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
HCQA @ Ego4D EgoSchema Challenge 2024
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
ObjectNLQ @ Ego4D Episodic Memory Challenge 2024
von: Feng, Yisen, et al.
Veröffentlicht: (2024)
von: Feng, Yisen, et al.
Veröffentlicht: (2024)
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
von: Hu, Yupeng, et al.
Veröffentlicht: (2023)
von: Hu, Yupeng, et al.
Veröffentlicht: (2023)
Detecting Deepfakes via Hamiltonian Dynamics
von: Cheng, Harry, et al.
Veröffentlicht: (2026)
von: Cheng, Harry, et al.
Veröffentlicht: (2026)
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
von: Feng, Yisen, et al.
Veröffentlicht: (2026)
von: Feng, Yisen, et al.
Veröffentlicht: (2026)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
Uncovering Hidden Subspaces in Video Diffusion Models Using Re-Identification
von: Dombrowski, Mischa, et al.
Veröffentlicht: (2024)
von: Dombrowski, Mischa, et al.
Veröffentlicht: (2024)
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
von: Lyu, Yibo, et al.
Veröffentlicht: (2026)
von: Lyu, Yibo, et al.
Veröffentlicht: (2026)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
R^3-VQA: "Read the Room" by Video Social Reasoning
von: Niu, Lixing, et al.
Veröffentlicht: (2025)
von: Niu, Lixing, et al.
Veröffentlicht: (2025)
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
Transition Matching Distillation for Fast Video Generation
von: Nie, Weili, et al.
Veröffentlicht: (2026)
von: Nie, Weili, et al.
Veröffentlicht: (2026)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
von: Chen, Yandu, et al.
Veröffentlicht: (2025)
von: Chen, Yandu, et al.
Veröffentlicht: (2025)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
von: Liu, Yating, et al.
Veröffentlicht: (2023)
von: Liu, Yating, et al.
Veröffentlicht: (2023)
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
von: Du, Hang, et al.
Veröffentlicht: (2024)
von: Du, Hang, et al.
Veröffentlicht: (2024)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
One-step Diffusion Models with $f$-Divergence Distribution Matching
von: Xu, Yilun, et al.
Veröffentlicht: (2025)
von: Xu, Yilun, et al.
Veröffentlicht: (2025)
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
von: Song, Enxin, et al.
Veröffentlicht: (2025)
von: Song, Enxin, et al.
Veröffentlicht: (2025)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Diffusion Facial Forgery Detection
von: Cheng, Harry, et al.
Veröffentlicht: (2024)
von: Cheng, Harry, et al.
Veröffentlicht: (2024)
MaeFuse: Transferring Omni Features with Pretrained Masked Autoencoders for Infrared and Visible Image Fusion via Guided Training
von: Li, Jiayang, et al.
Veröffentlicht: (2024)
von: Li, Jiayang, et al.
Veröffentlicht: (2024)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
von: Pang, Wei, et al.
Veröffentlicht: (2024)
von: Pang, Wei, et al.
Veröffentlicht: (2024)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
CoV: Chain-of-View Prompting for Spatial Reasoning
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025) -
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026) -
OSGNet @ Ego4D Episodic Memory Challenge 2025
von: Feng, Yisen, et al.
Veröffentlicht: (2025) -
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025) -
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)