Saved in:
| Main Authors: | Tamoyan, Hovhannes, Schuff, Hendrik, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.03974 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
by: Tamoyan, Hovhannes, et al.
Published: (2025)
by: Tamoyan, Hovhannes, et al.
Published: (2025)
How are Prompts Different in Terms of Sensitivity?
by: Lu, Sheng, et al.
Published: (2023)
by: Lu, Sheng, et al.
Published: (2023)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
by: Beck, Tilman, et al.
Published: (2023)
by: Beck, Tilman, et al.
Published: (2023)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
by: Fang, Haishuo, et al.
Published: (2024)
by: Fang, Haishuo, et al.
Published: (2024)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)
by: Dycke, Nils, et al.
Published: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
by: Bates, Luke, et al.
Published: (2023)
by: Bates, Luke, et al.
Published: (2023)
Citation Failure: Definition, Analysis and Efficient Mitigation
by: Buchmann, Jan, et al.
Published: (2025)
by: Buchmann, Jan, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
Commitment Checklist: Auditing Author Commitments in Peer Review
by: Chen, Chung-Chi, et al.
Published: (2026)
by: Chen, Chung-Chi, et al.
Published: (2026)
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
by: Afzal, Osama Mohammed, et al.
Published: (2025)
by: Afzal, Osama Mohammed, et al.
Published: (2025)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)
by: Paul, Indraneil, et al.
Published: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs
by: Fang, Haishuo, et al.
Published: (2024)
by: Fang, Haishuo, et al.
Published: (2024)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
by: Falk, Neele, et al.
Published: (2024)
by: Falk, Neele, et al.
Published: (2024)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
by: Ruan, Qian, et al.
Published: (2024)
by: Ruan, Qian, et al.
Published: (2024)
Attribute or Abstain: Large Language Models as Long Document Assistants
by: Buchmann, Jan, et al.
Published: (2024)
by: Buchmann, Jan, et al.
Published: (2024)
Identifying Aspects in Peer Reviews
by: Lu, Sheng, et al.
Published: (2025)
by: Lu, Sheng, et al.
Published: (2025)
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
by: Waldis, Andreas, et al.
Published: (2023)
by: Waldis, Andreas, et al.
Published: (2023)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
COVE: COntext and VEracity prediction for out-of-context images
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
Expert Preference-based Evaluation of Automated Related Work Generation
by: Şahinuç, Furkan, et al.
Published: (2025)
by: Şahinuç, Furkan, et al.
Published: (2025)
M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
by: Sachdeva, Rachneet, et al.
Published: (2023)
by: Sachdeva, Rachneet, et al.
Published: (2023)
Reward Modeling for Scientific Writing Evaluation
by: Şahinuç, Furkan, et al.
Published: (2026)
by: Şahinuç, Furkan, et al.
Published: (2026)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
by: Schuff, Hendrik, et al.
Published: (2021)
by: Schuff, Hendrik, et al.
Published: (2021)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
by: Purkayastha, Sukannya, et al.
Published: (2026)
by: Purkayastha, Sukannya, et al.
Published: (2026)
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
by: Sahnan, Dhruv, et al.
Published: (2026)
by: Sahnan, Dhruv, et al.
Published: (2026)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
by: Rohweder, Jonas, et al.
Published: (2026)
by: Rohweder, Jonas, et al.
Published: (2026)
DAPR: A Benchmark on Document-Aware Passage Retrieval
by: Wang, Kexin, et al.
Published: (2023)
by: Wang, Kexin, et al.
Published: (2023)
"Image, Tell me your story!" Predicting the original meta-context of visual misinformation
by: Tonglet, Jonathan, et al.
Published: (2024)
by: Tonglet, Jonathan, et al.
Published: (2024)
Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art
by: Liu, Chen Cecilia, et al.
Published: (2024)
by: Liu, Chen Cecilia, et al.
Published: (2024)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
by: Bigoulaeva, Irina, et al.
Published: (2025)
by: Bigoulaeva, Irina, et al.
Published: (2025)
Cultural Learning-Based Culture Adaptation of Language Models
by: Liu, Chen Cecilia, et al.
Published: (2025)
by: Liu, Chen Cecilia, et al.
Published: (2025)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
Similar Items
-
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
by: Tamoyan, Hovhannes, et al.
Published: (2025) -
How are Prompts Different in Terms of Sensitivity?
by: Lu, Sheng, et al.
Published: (2023) -
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
by: Beck, Tilman, et al.
Published: (2023) -
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
by: Fang, Haishuo, et al.
Published: (2024) -
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
by: Dycke, Nils, et al.
Published: (2025)