TVWorld: Foundations for Remote-Control TV Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Zhantao, Lu, Quanfeng, Zhong, Shuai, Yu, Dahai, Luo, Ping, Ng, Michael K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
von: Lu, Quanfeng, et al.
Veröffentlicht: (2025)
von: Lu, Quanfeng, et al.
Veröffentlicht: (2025)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
ADAM: An Embodied Causal Agent in Open-World Environments
von: Yu, Shu, et al.
Veröffentlicht: (2024)
von: Yu, Shu, et al.
Veröffentlicht: (2024)
TV2TV: A Unified Framework for Interleaved Language and Video Generation
von: Han, Xiaochuang, et al.
Veröffentlicht: (2025)
von: Han, Xiaochuang, et al.
Veröffentlicht: (2025)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
von: Jin, Haibo, et al.
Veröffentlicht: (2025)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
Efficient Adaptation For Remote Sensing Visual Grounding
von: Moughnieh, Hasan, et al.
Veröffentlicht: (2025)
von: Moughnieh, Hasan, et al.
Veröffentlicht: (2025)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
Interfacing Foundation Models' Embeddings
von: Zou, Xueyan, et al.
Veröffentlicht: (2023)
von: Zou, Xueyan, et al.
Veröffentlicht: (2023)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
von: Liang, Xiwen, et al.
Veröffentlicht: (2023)
von: Liang, Xiwen, et al.
Veröffentlicht: (2023)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
von: LASA Team, et al.
Veröffentlicht: (2025)
von: LASA Team, et al.
Veröffentlicht: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
von: Liang, Jiafeng, et al.
Veröffentlicht: (2025)
von: Liang, Jiafeng, et al.
Veröffentlicht: (2025)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
von: InternAgent Team, et al.
Veröffentlicht: (2025)
von: InternAgent Team, et al.
Veröffentlicht: (2025)
DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
von: Huang, Yidong, et al.
Veröffentlicht: (2024)
von: Huang, Yidong, et al.
Veröffentlicht: (2024)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
von: Su, Junhao, et al.
Veröffentlicht: (2025)
von: Su, Junhao, et al.
Veröffentlicht: (2025)
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
von: Ji, Yuxiang, et al.
Veröffentlicht: (2026)
von: Ji, Yuxiang, et al.
Veröffentlicht: (2026)
A Survey of Reasoning with Foundation Models
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
EPEE: Towards Efficient and Effective Foundation Models in Biomedicine
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
von: Ahn, Michael, et al.
Veröffentlicht: (2024)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2024)
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2024)
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
von: Lu, Quanfeng, et al.
Veröffentlicht: (2025) -
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024) -
ADAM: An Embodied Causal Agent in Open-World Environments
von: Yu, Shu, et al.
Veröffentlicht: (2024) -
TV2TV: A Unified Framework for Interleaved Language and Video Generation
von: Han, Xiaochuang, et al.
Veröffentlicht: (2025) -
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
von: Luo, Run, et al.
Veröffentlicht: (2025)