SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Jitesh, Li, Jialuo, Ma, Zixian, Zhang, Jieyu, Kim, Chris Dongjoo, Lee, Sangho, Tripathi, Rohun, Gupta, Tanmay, Clark, Christopher, Shi, Humphrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
VideoSAGE: Video Summarization with Graph Representation Learning
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
by: Chaves, Jose M. Rojas, et al.
Published: (2024)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
SAGE: Smart home Agent with Grounded Execution
by: Rivkin, Dmitriy, et al.
Published: (2023)
by: Rivkin, Dmitriy, et al.
Published: (2023)
Benchmarking Object Detectors with COCO: A New Path Forward
by: Singh, Shweta, et al.
Published: (2024)
by: Singh, Shweta, et al.
Published: (2024)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
SAGE: Multi-Agent Self-Evolution for LLM Reasoning
by: Peng, Yulin, et al.
Published: (2026)
by: Peng, Yulin, et al.
Published: (2026)
Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
Video-Based Reward Modeling for Computer-Use Agents
by: Song, Linxin, et al.
Published: (2026)
by: Song, Linxin, et al.
Published: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Reinforcement Learning for Long-Horizon Interactive LLM Agents
by: Chen, Kevin, et al.
Published: (2025)
by: Chen, Kevin, et al.
Published: (2025)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
PAI-Bench: A Comprehensive Benchmark For Physical AI
by: Zhou, Fengzhe, et al.
Published: (2025)
by: Zhou, Fengzhe, et al.
Published: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
The Case against Scale: Empirical Evidence of Underperformance in Large Secondary Funds
by: Gurav, Jitesh
Published: (2025)
by: Gurav, Jitesh
Published: (2025)
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
by: Kalyan, Vivek, et al.
Published: (2025)
by: Kalyan, Vivek, et al.
Published: (2025)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent
by: Hu, Yuyang, et al.
Published: (2026)
by: Hu, Yuyang, et al.
Published: (2026)
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
by: Hu, Yuyang, et al.
Published: (2026)
by: Hu, Yuyang, et al.
Published: (2026)
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
by: Yan, Sikuan, et al.
Published: (2026)
by: Yan, Sikuan, et al.
Published: (2026)
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
by: Chandwani, Abhishek, et al.
Published: (2026)
by: Chandwani, Abhishek, et al.
Published: (2026)
Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
VeruSAGE: A Study of Agent-Based Verification for Rust Systems
by: Yang, Chenyuan, et al.
Published: (2025)
by: Yang, Chenyuan, et al.
Published: (2025)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
by: Zhang, Ruiyang, et al.
Published: (2026)
by: Zhang, Ruiyang, et al.
Published: (2026)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
by: Ni, Tianwei, et al.
Published: (2025)
by: Ni, Tianwei, et al.
Published: (2025)
The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents
by: Sah, Tanmay, et al.
Published: (2026)
by: Sah, Tanmay, et al.
Published: (2026)
FedAA: A Reinforcement Learning Perspective on Adaptive Aggregation for Fair and Robust Federated Learning
by: He, Jialuo, et al.
Published: (2024)
by: He, Jialuo, et al.
Published: (2024)
WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning
by: Yu, Xinmiao, et al.
Published: (2026)
by: Yu, Xinmiao, et al.
Published: (2026)
Similar Items
-
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026) -
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026) -
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026) -
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024) -
AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
by: Jain, Jitesh, et al.
Published: (2025)