Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Bang, Ma, Ruotian, Jiang, Qingxuan, Wang, Peisong, Chen, Jiaqi, Xie, Zheng, Chen, Xingyu, Wang, Yue, Ye, Fanghua, Li, Jian, Yang, Yifan, Tu, Zhaopeng, Li, Xiaolong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
von: Wang, Peisong, et al.
Veröffentlicht: (2025)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
von: Yi, Zihao, et al.
Veröffentlicht: (2025)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
von: Wang, Yue, et al.
Veröffentlicht: (2025)
von: Wang, Yue, et al.
Veröffentlicht: (2025)
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
von: Ma, Xinbei, et al.
Veröffentlicht: (2025)
von: Ma, Xinbei, et al.
Veröffentlicht: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
von: Yang, Ruihan, et al.
Veröffentlicht: (2026)
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
Sentient House: Designing for Discourse
von: Collins, Robert
Veröffentlicht: (2024)
von: Collins, Robert
Veröffentlicht: (2024)
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
Benchmarking LLMs via Uncertainty Quantification
von: Ye, Fanghua, et al.
Veröffentlicht: (2024)
von: Ye, Fanghua, et al.
Veröffentlicht: (2024)
The Societal Response to Potentially Sentient AI
von: Caviola, Lucius
Veröffentlicht: (2025)
von: Caviola, Lucius
Veröffentlicht: (2025)
CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming
von: Wang, Peisong, et al.
Veröffentlicht: (2026)
von: Wang, Peisong, et al.
Veröffentlicht: (2026)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Cross-modality Data Augmentation for End-to-End Sign Language Translation
von: Ye, Jinhui, et al.
Veröffentlicht: (2023)
von: Ye, Jinhui, et al.
Veröffentlicht: (2023)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
von: Wei, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wei, Xiaolong, et al.
Veröffentlicht: (2025)
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
von: Li, Zihan, et al.
Veröffentlicht: (2026)
von: Li, Zihan, et al.
Veröffentlicht: (2026)
RaSA: Rank-Sharing Low-Rank Adaptation
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
CoAct: A Global-Local Hierarchy for Autonomous Agent Collaboration
von: Hou, Xinming, et al.
Veröffentlicht: (2024)
von: Hou, Xinming, et al.
Veröffentlicht: (2024)
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
CA+: Cognition Augmented Counselor Agent Framework for Long-term Dynamic Client Engagement
von: Tang, Yuanrong, et al.
Veröffentlicht: (2025)
von: Tang, Yuanrong, et al.
Veröffentlicht: (2025)
Remember You: Understanding How Users Use Deadbots to Reconstruct Memories of the Deceased
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
The Deep Learning model of Higher-Lower-Order Cognition, Memory, and Affection- More General Than KAN
von: Tao, Jun-Bo, et al.
Veröffentlicht: (2022)
von: Tao, Jun-Bo, et al.
Veröffentlicht: (2022)
Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking
von: Xu, Nan, et al.
Veröffentlicht: (2023)
von: Xu, Nan, et al.
Veröffentlicht: (2023)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
von: Li, Yansi, et al.
Veröffentlicht: (2025)
von: Li, Yansi, et al.
Veröffentlicht: (2025)
On the Information Redundancy in Non-Autoregressive Translation
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
Nexus: Higher-Order Attention Mechanisms in Transformers
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
von: Wang, Yue, et al.
Veröffentlicht: (2025)
von: Wang, Yue, et al.
Veröffentlicht: (2025)
Agent-as-a-Judge
von: You, Runyang, et al.
Veröffentlicht: (2026)
von: You, Runyang, et al.
Veröffentlicht: (2026)
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
von: Liang, Tian, et al.
Veröffentlicht: (2023)
von: Liang, Tian, et al.
Veröffentlicht: (2023)
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
von: Pei, Jiahuan, et al.
Veröffentlicht: (2025)
von: Pei, Jiahuan, et al.
Veröffentlicht: (2025)
"Mapping What I Feel": Understanding Affective Geovisualization Design Through the Lens of People-Place Relationships
von: Lan, Xingyu, et al.
Veröffentlicht: (2025)
von: Lan, Xingyu, et al.
Veröffentlicht: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
von: Zhang, Zijing, et al.
Veröffentlicht: (2025)
Block Rotation is All You Need for MXFP4 Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
von: Wang, Peisong, et al.
Veröffentlicht: (2025) -
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
von: Yi, Zihao, et al.
Veröffentlicht: (2025) -
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025) -
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
von: Liu, Cheng, et al.
Veröffentlicht: (2025) -
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
von: Wang, Yue, et al.
Veröffentlicht: (2025)