SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Dongwei, Zhang, Jingyu, Weller, Orion, Weir, Nathaniel, Van Durme, Benjamin, Khashabi, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
por: Weller, Orion, et al.
Publicado: (2023)
por: Weller, Orion, et al.
Publicado: (2023)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
por: Weir, Nathaniel, et al.
Publicado: (2024)
por: Weir, Nathaniel, et al.
Publicado: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
por: Jiang, Dongwei, et al.
Publicado: (2024)
por: Jiang, Dongwei, et al.
Publicado: (2024)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
por: Zhang, Jingyu, et al.
Publicado: (2024)
por: Zhang, Jingyu, et al.
Publicado: (2024)
Learning to Reason via Program Generation, Emulation, and Search
por: Weir, Nathaniel, et al.
Publicado: (2024)
por: Weir, Nathaniel, et al.
Publicado: (2024)
Many-Tier Instruction Hierarchy in LLM Agents
por: Zhang, Jingyu, et al.
Publicado: (2026)
por: Zhang, Jingyu, et al.
Publicado: (2026)
Reframing Tax Law Entailment as Analogical Reasoning
por: Zou, Xinrui, et al.
Publicado: (2024)
por: Zou, Xinrui, et al.
Publicado: (2024)
Crystal: Characterizing Relative Impact of Scholarly Publications
por: Collison, Hannah, et al.
Publicado: (2026)
por: Collison, Hannah, et al.
Publicado: (2026)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
por: Sanders, Kate, et al.
Publicado: (2024)
por: Sanders, Kate, et al.
Publicado: (2024)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
Defending Against Disinformation Attacks in Open-Domain Question Answering
por: Weller, Orion, et al.
Publicado: (2022)
por: Weller, Orion, et al.
Publicado: (2022)
SEQR: Secure and Efficient QR-based LoRA Routing
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
SpectR: Dynamically Composing LM Experts with Spectral Routing
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets
por: Weller, Orion, et al.
Publicado: (2023)
por: Weller, Orion, et al.
Publicado: (2023)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
por: Wang, Boshi, et al.
Publicado: (2024)
por: Wang, Boshi, et al.
Publicado: (2024)
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
por: Wang, Liaoyaqi, et al.
Publicado: (2025)
por: Wang, Liaoyaqi, et al.
Publicado: (2025)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
por: Wang, Hexuan, et al.
Publicado: (2026)
por: Wang, Hexuan, et al.
Publicado: (2026)
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
por: Jiang, Dongwei, et al.
Publicado: (2025)
por: Jiang, Dongwei, et al.
Publicado: (2025)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
por: Ahn, Jihyun Janice, et al.
Publicado: (2024)
por: Ahn, Jihyun Janice, et al.
Publicado: (2024)
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
por: Chari, Vivek, et al.
Publicado: (2025)
por: Chari, Vivek, et al.
Publicado: (2025)
SELF: Self-Evolution with Language Feedback
por: Lu, Jianqiao, et al.
Publicado: (2023)
por: Lu, Jianqiao, et al.
Publicado: (2023)
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
por: Cheng, Jeffrey, et al.
Publicado: (2024)
por: Cheng, Jeffrey, et al.
Publicado: (2024)
CORRECT: Context- and Reference-Augmented Reasoning and Prompting for Fact-Checking
por: Zhang, Delvin Ce, et al.
Publicado: (2025)
por: Zhang, Delvin Ce, et al.
Publicado: (2025)
Do Androids Know They're Only Dreaming of Electric Sheep?
por: CH-Wang, Sky, et al.
Publicado: (2023)
por: CH-Wang, Sky, et al.
Publicado: (2023)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
por: Hu, Michael Y., et al.
Publicado: (2025)
por: Hu, Michael Y., et al.
Publicado: (2025)
Core: Robust Factual Precision with Informative Sub-Claim Identification
por: Jiang, Zhengping, et al.
Publicado: (2024)
por: Jiang, Zhengping, et al.
Publicado: (2024)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
por: Mishra, Aayush, et al.
Publicado: (2025)
por: Mishra, Aayush, et al.
Publicado: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
por: Shen, Lingfeng, et al.
Publicado: (2023)
por: Shen, Lingfeng, et al.
Publicado: (2023)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
por: Chari, Vivek, et al.
Publicado: (2025)
por: Chari, Vivek, et al.
Publicado: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
por: Malek, Alan, et al.
Publicado: (2025)
por: Malek, Alan, et al.
Publicado: (2025)
NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
por: Weir, Nathaniel, et al.
Publicado: (2022)
por: Weir, Nathaniel, et al.
Publicado: (2022)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
por: Corallo, Giulio, et al.
Publicado: (2025)
por: Corallo, Giulio, et al.
Publicado: (2025)
NevIR: Negation in Neural Information Retrieval
por: Weller, Orion, et al.
Publicado: (2023)
por: Weller, Orion, et al.
Publicado: (2023)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
por: Chen, Tong, et al.
Publicado: (2024)
por: Chen, Tong, et al.
Publicado: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
por: Subramani, Nishant, et al.
Publicado: (2025)
por: Subramani, Nishant, et al.
Publicado: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026)
por: Zhao, Minda, et al.
Publicado: (2026)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
por: Shen, Lingfeng, et al.
Publicado: (2024)
por: Shen, Lingfeng, et al.
Publicado: (2024)
Ejemplares similares
-
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
por: Weller, Orion, et al.
Publicado: (2023) -
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
por: Weir, Nathaniel, et al.
Publicado: (2024) -
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
por: Jiang, Dongwei, et al.
Publicado: (2024) -
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
por: Zhang, Jingyu, et al.
Publicado: (2024) -
Learning to Reason via Program Generation, Emulation, and Search
por: Weir, Nathaniel, et al.
Publicado: (2024)