Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yan, Weicai, Dai, Yuhong, Ran, Qi, Li, Haodong, Lin, Wang, Jin, Tao, Xie, Xing, Liao, Hao, Lian, Jianxun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025)
par: Qin, Jialong, et autres
Publié: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
Interactive AI NPCs Powered by LLMs: Technical Report for the CPDC Challenge 2025
par: Huang, Yitian, et autres
Publié: (2025)
par: Huang, Yitian, et autres
Publié: (2025)
ProactBench: Beyond What The User Asked For
par: Harfi, Sepehr, et autres
Publié: (2026)
par: Harfi, Sepehr, et autres
Publié: (2026)
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
par: Li, Hongyu, et autres
Publié: (2025)
par: Li, Hongyu, et autres
Publié: (2025)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
Lost in Time: A New Temporal Benchmark for VideoLLMs
par: Cores, Daniel, et autres
Publié: (2024)
par: Cores, Daniel, et autres
Publié: (2024)
HumanLLM: Towards Personalized Understanding and Simulation of Human Nature
par: Lei, Yuxuan, et autres
Publié: (2026)
par: Lei, Yuxuan, et autres
Publié: (2026)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
par: Wu, Shiwei, et autres
Publié: (2024)
par: Wu, Shiwei, et autres
Publié: (2024)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
par: Wang, Han, et autres
Publié: (2024)
par: Wang, Han, et autres
Publié: (2024)
Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations
par: Huang, Xu, et autres
Publié: (2023)
par: Huang, Xu, et autres
Publié: (2023)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
par: Liang, Yujia, et autres
Publié: (2025)
par: Liang, Yujia, et autres
Publié: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
par: Fang, Pengcheng, et autres
Publié: (2025)
par: Fang, Pengcheng, et autres
Publié: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
par: Wang, Haibo, et autres
Publié: (2024)
par: Wang, Haibo, et autres
Publié: (2024)
Contextualized Privacy Defense for LLM Agents
par: Wen, Yule, et autres
Publié: (2026)
par: Wen, Yule, et autres
Publié: (2026)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026)
par: Zhang, Yulin, et autres
Publié: (2026)
Ada-Retrieval: An Adaptive Multi-Round Retrieval Paradigm for Sequential Recommendations
par: Li, Lei, et autres
Publié: (2024)
par: Li, Lei, et autres
Publié: (2024)
Aligning Large Language Models for Controllable Recommendations
par: Lu, Wensheng, et autres
Publié: (2024)
par: Lu, Wensheng, et autres
Publié: (2024)
Geometry-Guided Camera Motion Understanding in VideoLLMs
par: Feng, Haoan, et autres
Publié: (2026)
par: Feng, Haoan, et autres
Publié: (2026)
RecAI: Leveraging Large Language Models for Next-Generation Recommender Systems
par: Lian, Jianxun, et autres
Publié: (2024)
par: Lian, Jianxun, et autres
Publié: (2024)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
par: Guan, Yiran, et autres
Publié: (2026)
par: Guan, Yiran, et autres
Publié: (2026)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
par: Chung, Hyungjin, et autres
Publié: (2025)
par: Chung, Hyungjin, et autres
Publié: (2025)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
par: Zhang, Ziyi, et autres
Publié: (2025)
par: Zhang, Ziyi, et autres
Publié: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
par: Kim, Minji, et autres
Publié: (2025)
par: Kim, Minji, et autres
Publié: (2025)
LLM-powered Multi-agent Framework for Goal-oriented Learning in Intelligent Tutoring System
par: Wang, Tianfu, et autres
Publié: (2025)
par: Wang, Tianfu, et autres
Publié: (2025)
Text-Guided Multi-Scale Frequency Representation Adaptation
par: Yan, Weicai, et autres
Publié: (2026)
par: Yan, Weicai, et autres
Publié: (2026)
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
par: Shi, Wenlong, et autres
Publié: (2026)
par: Shi, Wenlong, et autres
Publié: (2026)
RecExplainer: Aligning Large Language Models for Explaining Recommendation Models
par: Lei, Yuxuan, et autres
Publié: (2023)
par: Lei, Yuxuan, et autres
Publié: (2023)
Aligning Language Models for Versatile Text-based Item Retrieval
par: Lei, Yuxuan, et autres
Publié: (2024)
par: Lei, Yuxuan, et autres
Publié: (2024)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
par: Li, Jiameng, et autres
Publié: (2026)
par: Li, Jiameng, et autres
Publié: (2026)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
par: Cai, Jianfeng, et autres
Publié: (2025)
par: Cai, Jianfeng, et autres
Publié: (2025)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
par: Yong, Xixian, et autres
Publié: (2025)
par: Yong, Xixian, et autres
Publié: (2025)
Towards Practical Real-Time Neural Video Compression
par: Jia, Zhaoyang, et autres
Publié: (2025)
par: Jia, Zhaoyang, et autres
Publié: (2025)
PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time Retrieval
par: Niu, Zihan, et autres
Publié: (2025)
par: Niu, Zihan, et autres
Publié: (2025)
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
par: Gong, Nanxu, et autres
Publié: (2026)
par: Gong, Nanxu, et autres
Publié: (2026)
To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasks
par: Gong, Nanxu, et autres
Publié: (2026)
par: Gong, Nanxu, et autres
Publié: (2026)
Documents similaires
-
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025) -
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025) -
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026) -
Interactive AI NPCs Powered by LLMs: Technical Report for the CPDC Challenge 2025
par: Huang, Yitian, et autres
Publié: (2025) -
ProactBench: Beyond What The User Asked For
par: Harfi, Sepehr, et autres
Publié: (2026)