T*: Re-thinking Temporal Search for Long-Form Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Jinhui, Wang, Zihan, Sun, Haosen, Chandrasegaran, Keshigeyan, Durante, Zane, Eyzaguirre, Cristobal, Bisk, Yonatan, Niebles, Juan Carlos, Adeli, Ehsan, Fei-Fei, Li, Wu, Jiajun, Li, Manling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
Exploring Diffusion Transformer Designs via Grafting
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2025)
MindCube: Spatial Mental Modeling from Limited Views
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
Streaming Detection of Queried Event Start
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
Learning Model Successors
by: Chang, Yingshan, et al.
Published: (2025)
by: Chang, Yingshan, et al.
Published: (2025)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
EXPANSIÓN Y LÍMITES DE LA BUENA FE OBJETIVA – A PROPÓSITO DEL “PROYECTO DE PRINCIPIOS LATINOAMERICANOS DE DERECHO DE LOS CONTRATOS”
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
OccFusion: Rendering Occluded Humans with Generative Diffusion Priors
by: Sun, Adam, et al.
Published: (2024)
by: Sun, Adam, et al.
Published: (2024)
MIRAGE: The Illusion of Visual Understanding
by: Asadi, Mohammad, et al.
Published: (2026)
by: Asadi, Mohammad, et al.
Published: (2026)
Automated Physical Performance Battery as a Digital Marker for Alzheimer’s Disease and Mild Cognitive Impairment
by: Ehsan Adeli
Published: (2024)
by: Ehsan Adeli
Published: (2024)
Wild2Avatar: Rendering Humans Behind Occlusions
by: Xiang, Tiange, et al.
Published: (2023)
by: Xiang, Tiange, et al.
Published: (2023)
Model Inversion Robustness: Can Transfer Learning Help?
by: Ho, Sy-Tuyen, et al.
Published: (2024)
by: Ho, Sy-Tuyen, et al.
Published: (2024)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
by: Abdollahzadeh, Milad, et al.
Published: (2023)
by: Abdollahzadeh, Milad, et al.
Published: (2023)
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
by: Liu, Christina, et al.
Published: (2025)
by: Liu, Christina, et al.
Published: (2025)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Planning with the Views via Scene Self-Exploration
by: Wang, Kangrui, et al.
Published: (2026)
by: Wang, Kangrui, et al.
Published: (2026)
MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Revision
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
by: Jain, Vidhi, et al.
Published: (2024)
by: Jain, Vidhi, et al.
Published: (2024)
Gradient Localization Improves Lifelong Pretraining of Language Models
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Neuro-Symbolic Decoding of Neural Activity
by: Wang, Yanchen, et al.
Published: (2026)
by: Wang, Yanchen, et al.
Published: (2026)
Temporal Preference Optimization for Long-Form Video Understanding
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
HumanScore: Benchmarking Human Motions in Generated Videos
by: Fang, Yusu, et al.
Published: (2026)
by: Fang, Yusu, et al.
Published: (2026)
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
by: Huang, Wenlong, et al.
Published: (2024)
by: Huang, Wenlong, et al.
Published: (2024)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
by: Thompson, Jacob, et al.
Published: (2025)
by: Thompson, Jacob, et al.
Published: (2025)
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
by: Xiang, Tiange, et al.
Published: (2025)
by: Xiang, Tiange, et al.
Published: (2025)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026)
by: Baade, Alan, et al.
Published: (2026)
Mechanical Processing Aggravates Cycling Degradations of LiCoO 2
by: Jinhui Li, et al.
Published: (2025)
by: Jinhui Li, et al.
Published: (2025)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
Similar Items
-
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024) -
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026) -
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026) -
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026) -
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)