Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Chunlei, Feng, Pengbin, Fu, Rong, Lee, Hoi Leong, Du, Xiaojing, Kang, Zhaolu, Zhang, Zeyu, Zhou, Weilin, Ouyang, Chun, Gan, Zhongxue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
by: Zhang, Meishan, et al.
Published: (2024)
by: Zhang, Meishan, et al.
Published: (2024)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
by: Zhuang, Yan, et al.
Published: (2025)
by: Zhuang, Yan, et al.
Published: (2025)
ModalImmune: Immunity Driven Unlearning via Self Destructive Training
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
by: Li, Zongyi, et al.
Published: (2025)
by: Li, Zongyi, et al.
Published: (2025)
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
by: Chen, Xiaolin, et al.
Published: (2025)
by: Chen, Xiaolin, et al.
Published: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Differentially Processed Optimized Collaborative Rich Text Editor
by: Jatana, Nishtha, et al.
Published: (2024)
by: Jatana, Nishtha, et al.
Published: (2024)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
by: Liu, Ke, et al.
Published: (2025)
by: Liu, Ke, et al.
Published: (2025)
Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
by: Guo, Ziheng, et al.
Published: (2025)
by: Guo, Ziheng, et al.
Published: (2025)
Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok
by: Zha, Mingyue, et al.
Published: (2026)
by: Zha, Mingyue, et al.
Published: (2026)
CDIO: Cross-Domain Inference Optimization with Resource Preference Prediction for Edge-Cloud Collaboration
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
by: Xu, Xinmeng, et al.
Published: (2026)
by: Xu, Xinmeng, et al.
Published: (2026)
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better
by: Ge, Mengying, et al.
Published: (2024)
by: Ge, Mengying, et al.
Published: (2024)
MInD: Improving Multimodal Sentiment Analysis via Multimodal Information Disentanglement
by: Dai, Weichen, et al.
Published: (2024)
by: Dai, Weichen, et al.
Published: (2024)
Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency
by: Zarghani, Abolfazl, et al.
Published: (2025)
by: Zarghani, Abolfazl, et al.
Published: (2025)
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
by: Jin, Zeyu, et al.
Published: (2026)
by: Jin, Zeyu, et al.
Published: (2026)
Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied Models
by: Liu, Shuyuan, et al.
Published: (2024)
by: Liu, Shuyuan, et al.
Published: (2024)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
by: Farina, Matteo, et al.
Published: (2025)
by: Farina, Matteo, et al.
Published: (2025)
Digital Fingerprinting on Multimedia: A Survey
by: Chen, Wendi, et al.
Published: (2024)
by: Chen, Wendi, et al.
Published: (2024)
Two-stage dynamic creative optimization under sparse ambiguous samples for e-commerce advertising
by: Li, Guandong, et al.
Published: (2023)
by: Li, Guandong, et al.
Published: (2023)
An automatic mixing speech enhancement system for multi-track audio
by: Liu, Xiaojing, et al.
Published: (2024)
by: Liu, Xiaojing, et al.
Published: (2024)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
by: Ling, Zeyu, et al.
Published: (2025)
by: Ling, Zeyu, et al.
Published: (2025)
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
by: Chen, Yi-Chun
Published: (2025)
by: Chen, Yi-Chun
Published: (2025)
A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions
by: Xu, Jinfeng, et al.
Published: (2025)
by: Xu, Jinfeng, et al.
Published: (2025)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
by: Cao, Yuqin, et al.
Published: (2025)
by: Cao, Yuqin, et al.
Published: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Smaller is Better: Generative Models Can Power Short Video Preloading
by: Liu, Liming, et al.
Published: (2026)
by: Liu, Liming, et al.
Published: (2026)
Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning
by: Zhao, Zijian, et al.
Published: (2026)
by: Zhao, Zijian, et al.
Published: (2026)
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
by: Yin, Yongkang, et al.
Published: (2023)
by: Yin, Yongkang, et al.
Published: (2023)
Similar Items
-
Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026) -
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
by: Meng, Chunlei, et al.
Published: (2026) -
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026) -
Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026) -
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
by: Zhang, Meishan, et al.
Published: (2024)