ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Huanzhen, Zhou, Ziheng, Song, Jiaqi, He, Li, Lan, Yunshi, Wang, Yan, Zhang, Wenqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling
by: Wang, Huanzhen, et al.
Published: (2026)
by: Wang, Huanzhen, et al.
Published: (2026)
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory
by: Lin, Yuxuan, et al.
Published: (2025)
by: Lin, Yuxuan, et al.
Published: (2025)
D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
A Survey on Facial Expression Recognition of Static and Dynamic Emotions
by: Wang, Yan, et al.
Published: (2024)
by: Wang, Yan, et al.
Published: (2024)
All rivers run into the sea: Unified Modality Brain-like Emotional Central Mechanism
by: Mai, Xinji, et al.
Published: (2024)
by: Mai, Xinji, et al.
Published: (2024)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
A$^{3}$lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP
by: Tao, Zeng, et al.
Published: (2024)
by: Tao, Zeng, et al.
Published: (2024)
DORA: Dynamic Online Reinforcement Agent for Token Merging in Vision Transformers
by: He, Kaixuan, et al.
Published: (2026)
by: He, Kaixuan, et al.
Published: (2026)
Text Adversarial Attacks with Dynamic Outputs
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Deciphering Functions of Neurons in Vision-Language Models
by: Xu, Jiaqi, et al.
Published: (2025)
by: Xu, Jiaqi, et al.
Published: (2025)
From Perception to Action: An Interactive Benchmark for Vision Reasoning
by: Wu, Yuhao, et al.
Published: (2026)
by: Wu, Yuhao, et al.
Published: (2026)
EmoScene: A Dual-space Dataset for Controllable Affective Image Generation
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
Fake It till You Make It: Curricular Dynamic Forgery Augmentations towards General Deepfake Detection
by: Lin, Yuzhen, et al.
Published: (2024)
by: Lin, Yuzhen, et al.
Published: (2024)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Unified Personalized Reward Model for Vision Generation
by: Wang, Yibin, et al.
Published: (2026)
by: Wang, Yibin, et al.
Published: (2026)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
by: Wang, Guanqun, et al.
Published: (2024)
by: Wang, Guanqun, et al.
Published: (2024)
Boosting Vision-Language Models for Histopathology Classification: Predict all at once
by: Zanella, Maxime, et al.
Published: (2024)
by: Zanella, Maxime, et al.
Published: (2024)
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
by: Tan, Hao, et al.
Published: (2026)
by: Tan, Hao, et al.
Published: (2026)
OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
by: Liu, Yaoli, et al.
Published: (2025)
by: Liu, Yaoli, et al.
Published: (2025)
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
by: He, Junwen, et al.
Published: (2024)
by: He, Junwen, et al.
Published: (2024)
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
by: Wang, Xinshun, et al.
Published: (2026)
by: Wang, Xinshun, et al.
Published: (2026)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs
by: Zhou, Ziheng, et al.
Published: (2026)
by: Zhou, Ziheng, et al.
Published: (2026)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
by: Liu, Xin, et al.
Published: (2023)
by: Liu, Xin, et al.
Published: (2023)
Raw Data Matters: Enhancing Prompt Tuning by Internal Augmentation on Vision-Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training
by: Huang, Zizheng, et al.
Published: (2024)
by: Huang, Zizheng, et al.
Published: (2024)
Occluded Human Pose Estimation based on Limb Joint Augmentation
by: Han, Gangtao, et al.
Published: (2024)
by: Han, Gangtao, et al.
Published: (2024)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
by: Li, Xinhao, et al.
Published: (2025)
by: Li, Xinhao, et al.
Published: (2025)
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
by: Yan, Ziang, et al.
Published: (2025)
by: Yan, Ziang, et al.
Published: (2025)
Few-shot Adaptation of Medical Vision-Language Models
by: Shakeri, Fereshteh, et al.
Published: (2024)
by: Shakeri, Fereshteh, et al.
Published: (2024)
EgoEMG: A Multimodal Egocentric Dataset with Bilateral EMG and Vision for Hand Pose Estimation
by: Xi, Ziheng, et al.
Published: (2026)
by: Xi, Ziheng, et al.
Published: (2026)
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Vision-Based Deep Reinforcement Learning of UAV Autonomous Navigation Using Privileged Information
by: Wang, Junqiao, et al.
Published: (2024)
by: Wang, Junqiao, et al.
Published: (2024)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
Active Learning via Vision-Language Model Adaptation with Open Data
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Similar Items
-
Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling
by: Wang, Huanzhen, et al.
Published: (2026) -
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
by: Wang, Xiaolong, et al.
Published: (2025) -
Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory
by: Lin, Yuxuan, et al.
Published: (2025) -
D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition
by: Wang, Haoran, et al.
Published: (2024) -
A Survey on Facial Expression Recognition of Static and Dynamic Emotions
by: Wang, Yan, et al.
Published: (2024)