Shared Imagination: LLMs Hallucinate Alike
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yilun, Xiong, Caiming, Savarese, Silvio, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
von: Le, Hung, et al.
Veröffentlicht: (2024)
von: Le, Hung, et al.
Veröffentlicht: (2024)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
von: Li, Jierui, et al.
Veröffentlicht: (2024)
von: Li, Jierui, et al.
Veröffentlicht: (2024)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
Agentic Confidence Calibration
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
Unanswerability Evaluation for Retrieval Augmented Generation
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
Editing Arbitrary Propositions in LLMs without Subject Labels
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
xGen-small Technical Report
von: Nijkamp, Erik, et al.
Veröffentlicht: (2025)
von: Nijkamp, Erik, et al.
Veröffentlicht: (2025)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
BingoGuard: LLM Content Moderation Tools with Risk Levels
von: Yin, Fan, et al.
Veröffentlicht: (2025)
von: Yin, Fan, et al.
Veröffentlicht: (2025)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Entropy-Based Block Pruning for Efficient Large Language Models
von: Yang, Liangwei, et al.
Veröffentlicht: (2025)
von: Yang, Liangwei, et al.
Veröffentlicht: (2025)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2025)
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2025)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Agentic Uncertainty Quantification
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2025)
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2025)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI
von: Zhang, Jianguo, et al.
Veröffentlicht: (2023)
von: Zhang, Jianguo, et al.
Veröffentlicht: (2023)
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
von: Murthy, Rithesh, et al.
Veröffentlicht: (2025)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
SSR: Socratic Self-Refine for Large Language Model Reasoning
von: Shi, Haizhou, et al.
Veröffentlicht: (2025)
von: Shi, Haizhou, et al.
Veröffentlicht: (2025)
Benchmarking Deep Search over Heterogeneous Enterprise Data
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2025)
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2025)
Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
von: Li, Gengyang, et al.
Veröffentlicht: (2026)
Text2Data: Low-Resource Data Generation with Textual Control
von: Wang, Shiyu, et al.
Veröffentlicht: (2024)
von: Wang, Shiyu, et al.
Veröffentlicht: (2024)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
What Makes Two Language Models Think Alike?
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
von: Salle, Jeanne, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
von: Pang, Bo, et al.
Veröffentlicht: (2025) -
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
von: Laban, Philippe, et al.
Veröffentlicht: (2024) -
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
von: Laban, Philippe, et al.
Veröffentlicht: (2023) -
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
von: Le, Hung, et al.
Veröffentlicht: (2024) -
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
von: Li, Jierui, et al.
Veröffentlicht: (2024)