AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhen, Yang, Yuhang, Jiang, Yunxiang, Lu, Yuhuan, Lu, Haifeng, Lian, Zheng, Zeng, Runhao, Hu, Xiping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opening the Black Box: Preliminary Insights into Affective Modeling in Multimodal Foundation Models
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Understanding Emotional Body Expressions via Large Language Models
by: Lu, Haifeng, et al.
Published: (2024)
by: Lu, Haifeng, et al.
Published: (2024)
Emotion Recognition from Skeleton Data: A Comprehensive Survey
by: Lu, Haifeng, et al.
Published: (2025)
by: Lu, Haifeng, et al.
Published: (2025)
Towards Stable Cross-Domain Depression Recognition under Missing Modalities
by: Chen, Jiuyi, et al.
Published: (2025)
by: Chen, Jiuyi, et al.
Published: (2025)
OVG-HQ: Online Video Grounding with Hybrid-modal Queries
by: Zeng, Runhao, et al.
Published: (2025)
by: Zeng, Runhao, et al.
Published: (2025)
CausalAffect: Causal Discovery for Facial Affective Understanding
by: Hu, Guanyu, et al.
Published: (2025)
by: Hu, Guanyu, et al.
Published: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models
by: Peng, Ying, et al.
Published: (2025)
by: Peng, Ying, et al.
Published: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Solution for 8th Competition on Affective & Behavior Analysis in-the-wild
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
by: Zeng, Runhao, et al.
Published: (2025)
by: Zeng, Runhao, et al.
Published: (2025)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026)
by: Yin, Xinlei, et al.
Published: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning
by: Zeng, Runhao, et al.
Published: (2024)
by: Zeng, Runhao, et al.
Published: (2024)
Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models
by: Zhang, Jingrui, et al.
Published: (2026)
by: Zhang, Jingrui, et al.
Published: (2026)
Temporal Action Detection Model Compression by Progressive Block Drop
by: Chen, Xiaoyong, et al.
Published: (2025)
by: Chen, Xiaoyong, et al.
Published: (2025)
Nesterov-Accelerated Robust Federated Learning Over Byzantine Adversaries
by: Xu, Lihan, et al.
Published: (2025)
by: Xu, Lihan, et al.
Published: (2025)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
by: Zhang, Xiaoyi, et al.
Published: (2025)
by: Zhang, Xiaoyi, et al.
Published: (2025)
Agentic Very Long Video Understanding
by: Rege, Aniket, et al.
Published: (2026)
by: Rege, Aniket, et al.
Published: (2026)
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
An Efficient Streaming Video Understanding Framework with Agentic Control
by: Liu, Jinming, et al.
Published: (2026)
by: Liu, Jinming, et al.
Published: (2026)
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
by: Deng, Qi, et al.
Published: (2024)
by: Deng, Qi, et al.
Published: (2024)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Balanced Co-Clustering of Users and Items for Embedding Table Compression in Recommender Systems
by: Jiang, Runhao, et al.
Published: (2026)
by: Jiang, Runhao, et al.
Published: (2026)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
by: Jiang, Chenxi, et al.
Published: (2025)
by: Jiang, Chenxi, et al.
Published: (2025)
Vagueness and the Connectives
by: Holliday, Wesley H.
Published: (2024)
by: Holliday, Wesley H.
Published: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026)
by: Zhao, Yiming, et al.
Published: (2026)
Query Understanding in LLM-based Conversational Information Seeking
by: Yuan, Yifei, et al.
Published: (2025)
by: Yuan, Yifei, et al.
Published: (2025)
Batch Query Processing and Optimization for Agentic Workflows
by: Shen, Junyi, et al.
Published: (2025)
by: Shen, Junyi, et al.
Published: (2025)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026)
by: Lin, Kuanwei, et al.
Published: (2026)
Vagueness Markers in Italian
by: Ghezzi, Chiara
Published: (2023)
by: Ghezzi, Chiara
Published: (2023)
DeRelayL: Sustainable Decentralized Relay Learning
by: Duan, Haihan, et al.
Published: (2026)
by: Duan, Haihan, et al.
Published: (2026)
Similar Items
-
Opening the Black Box: Preliminary Insights into Affective Modeling in Multimodal Foundation Models
by: Zhang, Zhen, et al.
Published: (2026) -
Understanding Emotional Body Expressions via Large Language Models
by: Lu, Haifeng, et al.
Published: (2024) -
Emotion Recognition from Skeleton Data: A Comprehensive Survey
by: Lu, Haifeng, et al.
Published: (2025) -
Towards Stable Cross-Domain Depression Recognition under Missing Modalities
by: Chen, Jiuyi, et al.
Published: (2025) -
OVG-HQ: Online Video Grounding with Hybrid-modal Queries
by: Zeng, Runhao, et al.
Published: (2025)