Look, Remember and Reason: Grounded reasoning in videos with language models
Fuente:
arXiv
Saved in:
| Main Authors: | Bhattacharyya, Apratim, Panchal, Sunny, Lee, Mingu, Pourreza, Reza, Madan, Pulkit, Memisevic, Roland |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
by: Pourreza, Reza, et al.
Published: (2025)
by: Pourreza, Reza, et al.
Published: (2025)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
by: Bhattacharyya, Apratim, et al.
Published: (2025)
by: Bhattacharyya, Apratim, et al.
Published: (2025)
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
by: Panchal, Sunny, et al.
Published: (2024)
by: Panchal, Sunny, et al.
Published: (2024)
Enhancing Hallucination Detection through Noise Injection
by: Liu, Litian, et al.
Published: (2025)
by: Liu, Litian, et al.
Published: (2025)
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
by: Haresh, Sanjay, et al.
Published: (2024)
by: Haresh, Sanjay, et al.
Published: (2024)
On the "Induction Bias" in Sequence Models
by: Ebrahimi, M. Reza, et al.
Published: (2026)
by: Ebrahimi, M. Reza, et al.
Published: (2026)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
Vision-language models lag human performance on physical dynamics and intent reasoning
by: Gu, Tianjun, et al.
Published: (2026)
by: Gu, Tianjun, et al.
Published: (2026)
Squeeze-and-Remember Block
by: Cakaj, Rinor, et al.
Published: (2024)
by: Cakaj, Rinor, et al.
Published: (2024)
Remembering Transformer for Continual Learning
by: Sun, Yuwei, et al.
Published: (2024)
by: Sun, Yuwei, et al.
Published: (2024)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers
by: Ebrahimi, MohammadReza, et al.
Published: (2024)
by: Ebrahimi, MohammadReza, et al.
Published: (2024)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
by: Ziai, Amir, et al.
Published: (2024)
by: Ziai, Amir, et al.
Published: (2024)
AirLetters: An Open Video Dataset of Characters Drawn in the Air
by: Dagli, Rishit, et al.
Published: (2024)
by: Dagli, Rishit, et al.
Published: (2024)
Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
In-distribution adversarial attacks on object recognition models using gradient-free search
by: Madan, Spandan, et al.
Published: (2021)
by: Madan, Spandan, et al.
Published: (2021)
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
by: Zhang, Hailiang, et al.
Published: (2024)
by: Zhang, Hailiang, et al.
Published: (2024)
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025)
by: Buckley, Thomas A., et al.
Published: (2025)
Reasoning emerges from constrained inference manifolds in large language models
by: Ma, Yanbiao, et al.
Published: (2026)
by: Ma, Yanbiao, et al.
Published: (2026)
Remembering by Reconstructing: Domain Incremental Learning With Test-Time Training on Video Streams
by: Swinnen, Jonathan, et al.
Published: (2026)
by: Swinnen, Jonathan, et al.
Published: (2026)
Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding
by: Kang, Minseok, et al.
Published: (2025)
by: Kang, Minseok, et al.
Published: (2025)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
by: Chu, Xu, et al.
Published: (2025)
by: Chu, Xu, et al.
Published: (2025)
Video models are zero-shot learners and reasoners
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
B-SMALL: A Bayesian Neural Network approach to Sparse Model-Agnostic Meta-Learning
by: Madan, Anish, et al.
Published: (2021)
by: Madan, Anish, et al.
Published: (2021)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
HyperCLIP: Adapting Vision-Language models with Hypernetworks
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
DL-EWF: Deep Learning Empowering Women's Fashion with Grounded-Segment-Anything Segmentation for Body Shape Classification
by: Asghari, Fatemeh, et al.
Published: (2024)
by: Asghari, Fatemeh, et al.
Published: (2024)
Text2Graph VPR: A Text-to-Graph Expert System for Explainable Place Recognition in Changing Environments
by: Yousefzadeh, Saeideh, et al.
Published: (2025)
by: Yousefzadeh, Saeideh, et al.
Published: (2025)
PDE-Constrained Optimization for Neural Image Segmentation with Physics Priors
by: Poudel, Seema K., et al.
Published: (2026)
by: Poudel, Seema K., et al.
Published: (2026)
CC-GRMAS: A Multi-Agent Graph Neural System for Spatiotemporal Landslide Risk Assessment in High Mountain Asia
by: Panchal, Mihir, et al.
Published: (2025)
by: Panchal, Mihir, et al.
Published: (2025)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
by: Liang, Zichen, et al.
Published: (2025)
by: Liang, Zichen, et al.
Published: (2025)
See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning
by: Wei, Yuxi, et al.
Published: (2026)
by: Wei, Yuxi, et al.
Published: (2026)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
Federated Learning with a Single Shared Image
by: Soni, Sunny, et al.
Published: (2024)
by: Soni, Sunny, et al.
Published: (2024)
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)
by: Venkatraman, Siddarth, et al.
Published: (2024)
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
by: Long, Lin, et al.
Published: (2025)
by: Long, Lin, et al.
Published: (2025)
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)
by: Wadhawan, Rohan, et al.
Published: (2025)
Similar Items
-
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
by: Pourreza, Reza, et al.
Published: (2025) -
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
by: Bhattacharyya, Apratim, et al.
Published: (2025) -
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
by: Panchal, Sunny, et al.
Published: (2024) -
Enhancing Hallucination Detection through Noise Injection
by: Liu, Litian, et al.
Published: (2025) -
ClevrSkills: Compositional Language and Visual Reasoning in Robotics
by: Haresh, Sanjay, et al.
Published: (2024)