RORA: Robust Free-Text Rationale Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Zhengping, Lu, Yining, Chen, Hanjie, Khashabi, Daniel, Van Durme, Benjamin, Liu, Anqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025)
Core: Robust Factual Precision with Informative Sub-Claim Identification
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2025)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2025)
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
von: Ou, Jiefu, et al.
Veröffentlicht: (2024)
von: Ou, Jiefu, et al.
Veröffentlicht: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Certified Mitigation of Worst-Case LLM Copyright Infringement
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
A Closer Look at Claim Decomposition
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2023)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2023)
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
von: Walden, William, et al.
Veröffentlicht: (2025)
von: Walden, William, et al.
Veröffentlicht: (2025)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
Benchmarking Language Model Creativity: A Case Study on Code Generation
von: Lu, Yining, et al.
Veröffentlicht: (2024)
von: Lu, Yining, et al.
Veröffentlicht: (2024)
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
von: Fang, Zhouxiang, et al.
Veröffentlicht: (2025)
von: Fang, Zhouxiang, et al.
Veröffentlicht: (2025)
Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
von: Li, Tianjian, et al.
Veröffentlicht: (2023)
von: Li, Tianjian, et al.
Veröffentlicht: (2023)
Evaluating the Evaluators: Are readability metrics good measures of readability?
von: Cachola, Isabel, et al.
Veröffentlicht: (2025)
von: Cachola, Isabel, et al.
Veröffentlicht: (2025)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SocialNLI: A Dialogue-Centric Social Inference Dataset
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
von: Jurayj, William, et al.
Veröffentlicht: (2025)
von: Jurayj, William, et al.
Veröffentlicht: (2025)
NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
von: Weir, Nathaniel, et al.
Veröffentlicht: (2022)
von: Weir, Nathaniel, et al.
Veröffentlicht: (2022)
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
Jailbreak Distillation: Renewable Safety Benchmarking
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024) -
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
von: Jiang, Zhengping, et al.
Veröffentlicht: (2025) -
Core: Robust Factual Precision with Informative Sub-Claim Identification
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024) -
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2025) -
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)