Reframe Anything: LLM Agent for Open World Video Reframing
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Jiawang, Wu, Yongliang, Chi, Weiheng, Zhu, Wenbo, Su, Ziyue, Wu, Jay |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
by: Leng, Zikang, et al.
Published: (2025)
by: Leng, Zikang, et al.
Published: (2025)
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
by: Wu, Meiqi, et al.
Published: (2024)
by: Wu, Meiqi, et al.
Published: (2024)
OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
by: Duan, Junwen, et al.
Published: (2025)
by: Duan, Junwen, et al.
Published: (2025)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
by: Zheng, Ruizhe, et al.
Published: (2024)
by: Zheng, Ruizhe, et al.
Published: (2024)
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
by: He, Yuchen, et al.
Published: (2025)
by: He, Yuchen, et al.
Published: (2025)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
ChatStitch: Visualizing Through Structures via Surround-View Unsupervised Deep Image Stitching with Collaborative LLM-Agents
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Clutter Detection and Removal by Multi-Objective Analysis for Photographic Guidance
by: Wu, Xiaoran
Published: (2025)
by: Wu, Xiaoran
Published: (2025)
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
by: Zhu, Yitong, et al.
Published: (2025)
by: Zhu, Yitong, et al.
Published: (2025)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
by: He, Xu, et al.
Published: (2024)
by: He, Xu, et al.
Published: (2024)
VideoMix: Aggregating How-To Videos for Task-Oriented Learning
by: Yang, Saelyne, et al.
Published: (2025)
by: Yang, Saelyne, et al.
Published: (2025)
AI-Enhanced Virtual Reality in Medicine: A Comprehensive Survey
by: Wu, Yixuan, et al.
Published: (2024)
by: Wu, Yixuan, et al.
Published: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
by: Cheng, Kanzhi, et al.
Published: (2026)
by: Cheng, Kanzhi, et al.
Published: (2026)
Enhancing Saliency Prediction in Monitoring Tasks: The Role of Visual Highlights
by: Wu, Zekun, et al.
Published: (2024)
by: Wu, Zekun, et al.
Published: (2024)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
by: Mazumdar, Amrita, et al.
Published: (2026)
by: Mazumdar, Amrita, et al.
Published: (2026)
Zero-Shot Segmentation of Eye Features Using the Segment Anything Model (SAM)
by: Maquiling, Virmarie, et al.
Published: (2023)
by: Maquiling, Virmarie, et al.
Published: (2023)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
by: Xu, Yutong, et al.
Published: (2024)
by: Xu, Yutong, et al.
Published: (2024)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
by: Nadeem, Asmar, et al.
Published: (2024)
by: Nadeem, Asmar, et al.
Published: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
by: Huh, Mina, et al.
Published: (2025)
by: Huh, Mina, et al.
Published: (2025)
Analyzing Swimming Performance Using Drone Captured Aerial Videos
by: Tran, Thu, et al.
Published: (2025)
by: Tran, Thu, et al.
Published: (2025)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
by: Eing, Lennart, et al.
Published: (2026)
by: Eing, Lennart, et al.
Published: (2026)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
Between Puppet and Actor: Reframing Authorship in this Age of AI Agents
by: Sun, Yuqian, et al.
Published: (2025)
by: Sun, Yuqian, et al.
Published: (2025)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2024)
by: Garg, Mallika, et al.
Published: (2024)
Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
by: Zhou, Puqi, et al.
Published: (2026)
by: Zhou, Puqi, et al.
Published: (2026)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
YOLOA: Real-Time Affordance Detection via LLM Adapter
by: Ji, Yuqi, et al.
Published: (2025)
by: Ji, Yuqi, et al.
Published: (2025)
VisionCAD: An Integration-Free Radiology Copilot Framework
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
by: Bao, Yiming, et al.
Published: (2024)
by: Bao, Yiming, et al.
Published: (2024)
egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks
by: Jammot, Matthias, et al.
Published: (2025)
by: Jammot, Matthias, et al.
Published: (2025)
DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition
by: Zhang, Ren, et al.
Published: (2025)
by: Zhang, Ren, et al.
Published: (2025)
CinePreGen: Camera Controllable Video Previsualization via Engine-powered Diffusion
by: Chen, Yiran, et al.
Published: (2024)
by: Chen, Yiran, et al.
Published: (2024)
Similar Items
-
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024) -
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
by: Wu, Yongliang, et al.
Published: (2024) -
AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
by: Leng, Zikang, et al.
Published: (2025) -
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025) -
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
by: Wu, Meiqi, et al.
Published: (2024)