The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Dongxin, Wu, Jikun, Yiu, Siu Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Model Collapse as Cultural Evolution
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation
von: Wu, Jikun, et al.
Veröffentlicht: (2026)
von: Wu, Jikun, et al.
Veröffentlicht: (2026)
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
von: Guo, Dongxin
Veröffentlicht: (2026)
von: Guo, Dongxin
Veröffentlicht: (2026)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
von: Vera, Patricio
Veröffentlicht: (2026)
von: Vera, Patricio
Veröffentlicht: (2026)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
In-Context Fixation: When Demonstrated Labels Override Semantics in Few-Shot Classification
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
von: Shafique, Muhammad Ali, et al.
Veröffentlicht: (2026)
von: Shafique, Muhammad Ali, et al.
Veröffentlicht: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
von: Redkar, Neel
Veröffentlicht: (2024)
von: Redkar, Neel
Veröffentlicht: (2024)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
von: Patarlapalli, Sai Babu, et al.
Veröffentlicht: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law
von: Shportun, Nazarii
Veröffentlicht: (2026)
von: Shportun, Nazarii
Veröffentlicht: (2026)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhilong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Model Collapse as Cultural Evolution
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation
von: Wu, Jikun, et al.
Veröffentlicht: (2026) -
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)