Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Caixin, Yan, Tianyu, Gong, Sitong, Zhang, Mingfang, Ouyang, Liangyang, Liu, Ruicong, Zheng, Bo, Lu, Huchuan, Zhang, Kaipeng, Sato, Yoichi, Huang, Yifei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2026)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2026)
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
von: Huang, Yifei, et al.
Veröffentlicht: (2025)
von: Huang, Yifei, et al.
Veröffentlicht: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
von: Liu, Ruicong, et al.
Veröffentlicht: (2024)
von: Liu, Ruicong, et al.
Veröffentlicht: (2024)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
von: Lin, Nie, et al.
Veröffentlicht: (2024)
von: Lin, Nie, et al.
Veröffentlicht: (2024)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
von: Lin, Nie, et al.
Veröffentlicht: (2025)
von: Lin, Nie, et al.
Veröffentlicht: (2025)
Towards Interactive Intelligence for Digital Humans
von: Cai, Yiyi, et al.
Veröffentlicht: (2025)
von: Cai, Yiyi, et al.
Veröffentlicht: (2025)
Parameter Aware Mamba Model for Multi-task Dense Prediction
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Enhancing Impression Change Prediction in Speed Dating Simulations Based on Speakers' Personalities
von: Matsuo, Kazuya, et al.
Veröffentlicht: (2025)
von: Matsuo, Kazuya, et al.
Veröffentlicht: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
Linking Perception, Confidence and Accuracy in MLLMs
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
von: Wang, Haowei, et al.
Veröffentlicht: (2023)
von: Wang, Haowei, et al.
Veröffentlicht: (2023)
Subjective Face Transform using Human First Impressions
von: Roygaga, Chaitanya, et al.
Veröffentlicht: (2023)
von: Roygaga, Chaitanya, et al.
Veröffentlicht: (2023)
Enhancing Representation Learning of EEG Data with Masked Autoencoders
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
Prompt and Prejudice
von: Berlincioni, Lorenzo, et al.
Veröffentlicht: (2024)
von: Berlincioni, Lorenzo, et al.
Veröffentlicht: (2024)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
Can Impressions of Music be Extracted from Thumbnail Images?
von: Harada, Takashi, et al.
Veröffentlicht: (2025)
von: Harada, Takashi, et al.
Veröffentlicht: (2025)
Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
Can MLLMs Perform Text-to-Image In-Context Learning?
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification
von: Zhang, Pingping, et al.
Veröffentlicht: (2025)
von: Zhang, Pingping, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
von: Sun, Chengbo, et al.
Veröffentlicht: (2025)
von: Sun, Chengbo, et al.
Veröffentlicht: (2025)
Explore the Hallucination on Low-level Perception for MLLMs
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025) -
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025) -
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2026) -
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
von: Liu, Ruicong, et al.
Veröffentlicht: (2025) -
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)