CAViAR: Critic-Augmented Video Agentic Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Menon, Sachit, Iscen, Ahmet, Nagrani, Arsha, Weyand, Tobias, Vondrick, Carl, Schmid, Cordelia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MINERVA: Evaluating Complex Video Reasoning
por: Nagrani, Arsha, et al.
Publicado: (2025)
por: Nagrani, Arsha, et al.
Publicado: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
por: Min, Juhong, et al.
Publicado: (2024)
por: Min, Juhong, et al.
Publicado: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2026)
por: Nagrani, Arsha, et al.
Publicado: (2026)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
por: Singh, Darshan, et al.
Publicado: (2026)
por: Singh, Darshan, et al.
Publicado: (2026)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2024)
por: Nagrani, Arsha, et al.
Publicado: (2024)
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
por: Menon, Sachit, et al.
Publicado: (2024)
por: Menon, Sachit, et al.
Publicado: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
por: Sokar, Ghada, et al.
Publicado: (2025)
por: Sokar, Ghada, et al.
Publicado: (2025)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
por: Caron, Mathilde, et al.
Publicado: (2024)
por: Caron, Mathilde, et al.
Publicado: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
por: Iscen, Ahmet, et al.
Publicado: (2023)
por: Iscen, Ahmet, et al.
Publicado: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
por: Caron, Mathilde, et al.
Publicado: (2024)
por: Caron, Mathilde, et al.
Publicado: (2024)
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)
por: Zhou, Xingyi, et al.
Publicado: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
por: Uijlings, Jasper, et al.
Publicado: (2025)
por: Uijlings, Jasper, et al.
Publicado: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
por: Ghosh, Partha, et al.
Publicado: (2024)
por: Ghosh, Partha, et al.
Publicado: (2024)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
por: Hu, Ziniu, et al.
Publicado: (2024)
por: Hu, Ziniu, et al.
Publicado: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
por: Kahatapitiya, Kumara, et al.
Publicado: (2023)
por: Kahatapitiya, Kumara, et al.
Publicado: (2023)
Generating Illustrated Instructions
por: Menon, Sachit, et al.
Publicado: (2023)
por: Menon, Sachit, et al.
Publicado: (2023)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
por: Kang, Dahyun, et al.
Publicado: (2025)
por: Kang, Dahyun, et al.
Publicado: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
por: Chen, Zerui, et al.
Publicado: (2024)
por: Chen, Zerui, et al.
Publicado: (2024)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
por: Jain, Gagan, et al.
Publicado: (2024)
por: Jain, Gagan, et al.
Publicado: (2024)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
por: Fiastre, Gabriel, et al.
Publicado: (2025)
por: Fiastre, Gabriel, et al.
Publicado: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
por: Shvetsova, Nina, et al.
Publicado: (2025)
por: Shvetsova, Nina, et al.
Publicado: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
por: Mercea, Otniel-Bogdan, et al.
Publicado: (2024)
por: Mercea, Otniel-Bogdan, et al.
Publicado: (2024)
DataDream: Few-shot Guided Dataset Generation
por: Kim, Jae Myung, et al.
Publicado: (2024)
por: Kim, Jae Myung, et al.
Publicado: (2024)
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
por: Arnab, Anurag, et al.
Publicado: (2025)
por: Arnab, Anurag, et al.
Publicado: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
por: Shen, Junhong, et al.
Publicado: (2025)
por: Shen, Junhong, et al.
Publicado: (2025)
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
por: Mall, Utkarsh, et al.
Publicado: (2025)
por: Mall, Utkarsh, et al.
Publicado: (2025)
pix2gestalt: Amodal Segmentation by Synthesizing Wholes
por: Ozguroglu, Ege, et al.
Publicado: (2024)
por: Ozguroglu, Ege, et al.
Publicado: (2024)
Do multimodal models imagine electric sheep?
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
Visual Lexicon: Rich Image Features in Language Space
por: Wang, XuDong, et al.
Publicado: (2024)
por: Wang, XuDong, et al.
Publicado: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
por: Kuhar, Sachit, et al.
Publicado: (2023)
por: Kuhar, Sachit, et al.
Publicado: (2023)
Agentic Very Long Video Understanding
por: Rege, Aniket, et al.
Publicado: (2026)
por: Rege, Aniket, et al.
Publicado: (2026)
Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2024)
por: Kazakos, Evangelos, et al.
Publicado: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2025)
por: Kazakos, Evangelos, et al.
Publicado: (2025)
GES: Generalized Exponential Splatting for Efficient Radiance Field Rendering
por: Hamdi, Abdullah, et al.
Publicado: (2024)
por: Hamdi, Abdullah, et al.
Publicado: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
por: Bousselham, Walid, et al.
Publicado: (2025)
por: Bousselham, Walid, et al.
Publicado: (2025)
AutoAD III: The Prequel -- Back to the Pixels
por: Han, Tengda, et al.
Publicado: (2024)
por: Han, Tengda, et al.
Publicado: (2024)
Language-Guided Image Tokenization for Generation
por: Zha, Kaiwen, et al.
Publicado: (2024)
por: Zha, Kaiwen, et al.
Publicado: (2024)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
por: Farid, Karim, et al.
Publicado: (2025)
por: Farid, Karim, et al.
Publicado: (2025)
How Video Meetings Change Your Expression
por: Sarin, Sumit, et al.
Publicado: (2024)
por: Sarin, Sumit, et al.
Publicado: (2024)
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
por: Zhang, Wenlun, et al.
Publicado: (2026)
por: Zhang, Wenlun, et al.
Publicado: (2026)
Ejemplares similares
-
MINERVA: Evaluating Complex Video Reasoning
por: Nagrani, Arsha, et al.
Publicado: (2025) -
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
por: Min, Juhong, et al.
Publicado: (2024) -
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2026) -
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
por: Singh, Darshan, et al.
Publicado: (2026) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2024)