GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Ming, Wu, Shuang, Wang, Bixuan, Lin, Lu, Chen, Yuxin, Yang, Xiaocui, Wang, Daling, Feng, Shi, Zhang, Yifei, Sun, Yufan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AnnaAgent: Dynamic Evolution Agent System with Multi-Session Memory for Realistic Seeker Simulation
por: Wang, Ming, et al.
Publicado: (2025)
por: Wang, Ming, et al.
Publicado: (2025)
Enhancing LLM-based Recommendation through Semantic-Aligned Collaborative Knowledge
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering
por: Zhang, Xiaoming, et al.
Publicado: (2024)
por: Zhang, Xiaoming, et al.
Publicado: (2024)
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
por: Zhao, Yinzhi, et al.
Publicado: (2026)
por: Zhao, Yinzhi, et al.
Publicado: (2026)
Language Models as Continuous Self-Evolving Data Engineers
por: Wang, Peidong, et al.
Publicado: (2024)
por: Wang, Peidong, et al.
Publicado: (2024)
TOOL-ED: Enhancing Empathetic Response Generation with the Tool Calling Capability of LLM
por: Cao, Huiying, et al.
Publicado: (2024)
por: Cao, Huiying, et al.
Publicado: (2024)
Generative Emotion Cause Explanation in Multimodal Conversations
por: Wang, Lin, et al.
Publicado: (2024)
por: Wang, Lin, et al.
Publicado: (2024)
From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning
por: Wang, Xiao, et al.
Publicado: (2026)
por: Wang, Xiao, et al.
Publicado: (2026)
Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
por: Wang, Zihan, et al.
Publicado: (2024)
por: Wang, Zihan, et al.
Publicado: (2024)
Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
por: Shangguan, Haonan, et al.
Publicado: (2025)
por: Shangguan, Haonan, et al.
Publicado: (2025)
GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
por: Wan, Faxian, et al.
Publicado: (2026)
por: Wan, Faxian, et al.
Publicado: (2026)
MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction
por: Mu, Shiyi, et al.
Publicado: (2025)
por: Mu, Shiyi, et al.
Publicado: (2025)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
por: Zhang, Yiqun, et al.
Publicado: (2024)
por: Zhang, Yiqun, et al.
Publicado: (2024)
Is Mamba Effective for Time Series Forecasting?
por: Wang, Zihan, et al.
Publicado: (2024)
por: Wang, Zihan, et al.
Publicado: (2024)
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
por: Gao, Zhuoyue, et al.
Publicado: (2026)
por: Gao, Zhuoyue, et al.
Publicado: (2026)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
por: Cai, Dexian, et al.
Publicado: (2025)
por: Cai, Dexian, et al.
Publicado: (2025)
MoLAN: A Unified Modality-Aware Noise Dynamic Editing Framework for Multimodal Sentiment Analysis
por: Xu, Xingle, et al.
Publicado: (2025)
por: Xu, Xingle, et al.
Publicado: (2025)
CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering
por: Wei, Zili, et al.
Publicado: (2026)
por: Wei, Zili, et al.
Publicado: (2026)
T-COL: Generating Counterfactual Explanations for General User Preferences on Variable Machine Learning Systems
por: Wang, Ming, et al.
Publicado: (2023)
por: Wang, Ming, et al.
Publicado: (2023)
NEAT: Neuron-Based Early Exit for Large Reasoning Models
por: Liu, Kang, et al.
Publicado: (2026)
por: Liu, Kang, et al.
Publicado: (2026)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
por: Yang, Xiaocui, et al.
Publicado: (2024)
por: Yang, Xiaocui, et al.
Publicado: (2024)
PsyDraw: A Multi-Agent Multimodal System for Mental Health Screening in Left-Behind Children
por: Zhang, Yiqun, et al.
Publicado: (2024)
por: Zhang, Yiqun, et al.
Publicado: (2024)
RoCar: A Relationship Network-based Evaluation Method for Large Language Models
por: Wang, Ming, et al.
Publicado: (2023)
por: Wang, Ming, et al.
Publicado: (2023)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
por: Zhang, Yiqun, et al.
Publicado: (2026)
por: Zhang, Yiqun, et al.
Publicado: (2026)
Minstrel: Structural Prompt Generation with Multi-Agents Coordination for Non-AI Experts
por: Wang, Ming, et al.
Publicado: (2024)
por: Wang, Ming, et al.
Publicado: (2024)
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
STICKERCONV: Generating Multimodal Empathetic Responses from Scratch
por: Zhang, Yiqun, et al.
Publicado: (2024)
por: Zhang, Yiqun, et al.
Publicado: (2024)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
por: Liu, Yongkang, et al.
Publicado: (2023)
por: Liu, Yongkang, et al.
Publicado: (2023)
ChatZero:Zero-shot Cross-Lingual Dialogue Generation via Pseudo-Target Language
por: Liu, Yongkang, et al.
Publicado: (2024)
por: Liu, Yongkang, et al.
Publicado: (2024)
Why Do More Experts Fail? A Theoretical Analysis of Model Merging
por: Wang, Zijing, et al.
Publicado: (2025)
por: Wang, Zijing, et al.
Publicado: (2025)
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
por: Wang, Zijing, et al.
Publicado: (2026)
por: Wang, Zijing, et al.
Publicado: (2026)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
por: Huang, YiJie, et al.
Publicado: (2026)
por: Huang, YiJie, et al.
Publicado: (2026)
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
por: Wang, Zijing, et al.
Publicado: (2026)
por: Wang, Zijing, et al.
Publicado: (2026)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
por: Kong, Fanheng, et al.
Publicado: (2026)
por: Kong, Fanheng, et al.
Publicado: (2026)
Affective Computing in the Era of Large Language Models: A Survey from the NLP Perspective
por: Zhang, Yiqun, et al.
Publicado: (2024)
por: Zhang, Yiqun, et al.
Publicado: (2024)
Identifiability of VAR(1) model in a stationary setting
por: Liu, Bixuan
Publicado: (2025)
por: Liu, Bixuan
Publicado: (2025)
SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
por: Wang, Peidong, et al.
Publicado: (2026)
por: Wang, Peidong, et al.
Publicado: (2026)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
por: Weng, Shihao, et al.
Publicado: (2026)
por: Weng, Shihao, et al.
Publicado: (2026)
Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
por: Liu, Yongkang, et al.
Publicado: (2025)
por: Liu, Yongkang, et al.
Publicado: (2025)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
por: Wu, Rong, et al.
Publicado: (2025)
por: Wu, Rong, et al.
Publicado: (2025)
Ejemplares similares
-
AnnaAgent: Dynamic Evolution Agent System with Multi-Session Memory for Realistic Seeker Simulation
por: Wang, Ming, et al.
Publicado: (2025) -
Enhancing LLM-based Recommendation through Semantic-Aligned Collaborative Knowledge
por: Wang, Zihan, et al.
Publicado: (2025) -
Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering
por: Zhang, Xiaoming, et al.
Publicado: (2024) -
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
por: Zhao, Yinzhi, et al.
Publicado: (2026) -
Language Models as Continuous Self-Evolving Data Engineers
por: Wang, Peidong, et al.
Publicado: (2024)