How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Rui, Xia, Heming, Yuan, Xinfeng, Dong, Qingxiu, Sha, Lei, Li, Wenjie, Sui, Zhifang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
von: Yang, Yixin, et al.
Veröffentlicht: (2026)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Towards Harmonized Uncertainty Estimation for Large Language Models
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Self-Boosting Large Language Models with Synthetic Preference Data
von: Dong, Qingxiu, et al.
Veröffentlicht: (2024)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2024)
Plug-and-Play Training Framework for Preference Optimization
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
Beyond Single Frames: Can LMMs Comprehend Temporal and Contextual Narratives in Image Sequences?
von: Wang, Xiaochen, et al.
Veröffentlicht: (2025)
von: Wang, Xiaochen, et al.
Veröffentlicht: (2025)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
A Survey on In-context Learning
von: Dong, Qingxiu, et al.
Veröffentlicht: (2022)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2022)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Digital Twins: How Far from Ideas to Twins?
von: Jingyu, Lu
Veröffentlicht: (2024)
von: Jingyu, Lu
Veröffentlicht: (2024)
Reinforcement Pre-Training
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025)
von: Du, Bangde, et al.
Veröffentlicht: (2025)
CoLT: Reasoning with Chain of Latent Tool Calls
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2026)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Chain-of-Thought Tokens are Computer Program Variables
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
von: Zhu, Fangwei, et al.
Veröffentlicht: (2025)
ISOMORPH: A Supply Chain Digital Twin for Simulation, Dataset Generation, and Forecasting Benchmarks
von: Zhang, Zhizhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhizhen, et al.
Veröffentlicht: (2026)
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
LLM-REVal: Can We Trust LLM Reviewers Yet?
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Enhancing Tool Retrieval with Iterative Feedback from Large Language Models
von: Xu, Qiancheng, et al.
Veröffentlicht: (2024)
von: Xu, Qiancheng, et al.
Veröffentlicht: (2024)
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens
von: Luo, Weiyao, et al.
Veröffentlicht: (2024)
von: Luo, Weiyao, et al.
Veröffentlicht: (2024)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Chain-of-Query: Unleashing the Power of LLMs in SQL-Aided Table Understanding via Multi-Agent Collaboration
von: Sui, Songyuan, et al.
Veröffentlicht: (2025)
von: Sui, Songyuan, et al.
Veröffentlicht: (2025)
How Far are App Secrets from Being Stolen? A Case Study on Android
von: Wei, Lili, et al.
Veröffentlicht: (2025)
von: Wei, Lili, et al.
Veröffentlicht: (2025)
The Digital Cybersecurity Expert: How Far Have We Come?
von: Wang, Dawei, et al.
Veröffentlicht: (2025)
von: Wang, Dawei, et al.
Veröffentlicht: (2025)
LLMs for Relational Reasoning: How Far are We?
von: Li, Zhiming, et al.
Veröffentlicht: (2024)
von: Li, Zhiming, et al.
Veröffentlicht: (2024)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
Interior Hessian estimates for Hessian quotient equations in dimension three
von: Jiao, Heming, et al.
Veröffentlicht: (2026)
von: Jiao, Heming, et al.
Veröffentlicht: (2026)
How Far Back in Time a Digital Twin Reflects the State of the Physical Object: Age of Staleness
von: Cosandal, Ismail, et al.
Veröffentlicht: (2026)
von: Cosandal, Ismail, et al.
Veröffentlicht: (2026)
Large Language Models Struggle with Unreasonability in Math Problems
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2024)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2026)
von: Xia, Heming, et al.
Veröffentlicht: (2026)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Character is Destiny: Can Role-Playing Language Agents Make Persona-Driven Decisions?
von: Xu, Rui, et al.
Veröffentlicht: (2024)
von: Xu, Rui, et al.
Veröffentlicht: (2024)
Model Editing for LLMs4Code: How Far are We?
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
von: Leong, Chak Tou, et al.
Veröffentlicht: (2026)
von: Leong, Chak Tou, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024) -
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2024) -
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
von: Yang, Yixin, et al.
Veröffentlicht: (2026) -
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025) -
Towards Harmonized Uncertainty Estimation for Large Language Models
von: Li, Rui, et al.
Veröffentlicht: (2025)