Gespeichert in:
| Hauptverfasser: | Sun, Xiangkun, Kong, Lingkai, Zhang, Aoqi, Zeng, Liang, Wang, Tonghan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.09314 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
von: Fesser, Lukas, et al.
Veröffentlicht: (2026)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
Empowering Tabular Data Preparation with Language Models: Why and How?
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
von: Chen, Mengshi, et al.
Veröffentlicht: (2025)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
von: Fang, Bowen, et al.
Veröffentlicht: (2026)
von: Fang, Bowen, et al.
Veröffentlicht: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
von: Peng, Songping, et al.
Veröffentlicht: (2026)
von: Peng, Songping, et al.
Veröffentlicht: (2026)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
von: Ivanov, Igor
Veröffentlicht: (2025)
von: Ivanov, Igor
Veröffentlicht: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2026)
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
von: Liu, Yawei
Veröffentlicht: (2025)
von: Liu, Yawei
Veröffentlicht: (2025)
Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
von: Bayarri-Planas, Jordi, et al.
Veröffentlicht: (2024)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
von: Rai, Daking, et al.
Veröffentlicht: (2024)
von: Rai, Daking, et al.
Veröffentlicht: (2024)
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
von: Davide, Fabrizio, et al.
Veröffentlicht: (2024)
von: Davide, Fabrizio, et al.
Veröffentlicht: (2024)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
von: Ansell, Rebecca, et al.
Veröffentlicht: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
von: Wang, Xiaohua, et al.
Veröffentlicht: (2026)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
von: Wang, Shuai, et al.
Veröffentlicht: (2021)
von: Wang, Shuai, et al.
Veröffentlicht: (2021)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
von: Iqbal, Hasan, et al.
Veröffentlicht: (2024)
von: Iqbal, Hasan, et al.
Veröffentlicht: (2024)
Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models
von: Liu, Ruitong, et al.
Veröffentlicht: (2025)
von: Liu, Ruitong, et al.
Veröffentlicht: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
von: Lai, Junyu, et al.
Veröffentlicht: (2025)
von: Lai, Junyu, et al.
Veröffentlicht: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
How Does Unfaithful Reasoning Emerge from Autoregressive Training? A Study of Synthetic Experiments
von: Wang, Fuxin, et al.
Veröffentlicht: (2026)
von: Wang, Fuxin, et al.
Veröffentlicht: (2026)
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions
von: Pan, Huitong, et al.
Veröffentlicht: (2024)
von: Pan, Huitong, et al.
Veröffentlicht: (2024)
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models
von: Darm, Paul, et al.
Veröffentlicht: (2025)
von: Darm, Paul, et al.
Veröffentlicht: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
von: Fang, Xi, et al.
Veröffentlicht: (2025)
von: Fang, Xi, et al.
Veröffentlicht: (2025)
CausalT5K: Diagnosing and Informing Refusal for Trustworthy Causal Reasoning of Skepticism, Sycophancy, Detection-Correction, and Rung Collapse
von: Geng, Longling, et al.
Veröffentlicht: (2026)
von: Geng, Longling, et al.
Veröffentlicht: (2026)
On Explaining with Attention Matrices
von: Naim, Omar, et al.
Veröffentlicht: (2024)
von: Naim, Omar, et al.
Veröffentlicht: (2024)
Process Supervision-Guided Policy Optimization for Code Generation
von: Dai, Ning, et al.
Veröffentlicht: (2024)
von: Dai, Ning, et al.
Veröffentlicht: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
von: He, Yongquan, et al.
Veröffentlicht: (2024)
von: He, Yongquan, et al.
Veröffentlicht: (2024)
Paying Attention to Deflections: Mining Pragmatic Nuances for Whataboutism Detection in Online Discourse
von: Phi, Khiem, et al.
Veröffentlicht: (2024)
von: Phi, Khiem, et al.
Veröffentlicht: (2024)
Large Language Models as 'Hidden Persuaders': Fake Product Reviews are Indistinguishable to Humans and Machines
von: Meng, Weiyao, et al.
Veröffentlicht: (2025)
von: Meng, Weiyao, et al.
Veröffentlicht: (2025)
Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific Discovery
von: Yuan, Xinzhe, et al.
Veröffentlicht: (2026)
von: Yuan, Xinzhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
von: Wang, Liang, et al.
Veröffentlicht: (2026) -
Evaluating Relational Reasoning in LLMs with REL
von: Fesser, Lukas, et al.
Veröffentlicht: (2026) -
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)