Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals
Fuente:
arXiv
Salvato in:
| Autori principali: | Elazar, Yanai, Paranjape, Bhargavi, Peng, Hao, Wiegreffe, Sarah, Raghavi, Khyathi, Srikumar, Vivek, Singh, Sameer, Smith, Noah A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Linear Representations and Pretraining Data Frequency in Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2025)
di: Merullo, Jack, et al.
Pubblicazione: (2025)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
di: Merrill, William, et al.
Pubblicazione: (2024)
di: Merrill, William, et al.
Pubblicazione: (2024)
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
di: Nadkarni, Rahul, et al.
Pubblicazione: (2025)
di: Nadkarni, Rahul, et al.
Pubblicazione: (2025)
Estimating the Causal Effect of Early ArXiving on Paper Acceptance
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
di: Elazar, Yanai, et al.
Pubblicazione: (2026)
di: Elazar, Yanai, et al.
Pubblicazione: (2026)
Detection and Measurement of Syntactic Templates in Generated Text
di: Shaib, Chantal, et al.
Pubblicazione: (2024)
di: Shaib, Chantal, et al.
Pubblicazione: (2024)
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
Reinforcing Code Generation: Improving Text-to-SQL with Execution-Based Learning
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
di: Iluz, Bar, et al.
Pubblicazione: (2024)
di: Iluz, Bar, et al.
Pubblicazione: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
What's In My Big Data?
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
di: Elazar, Yanai, et al.
Pubblicazione: (2023)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
di: Kim, Joongwon, et al.
Pubblicazione: (2024)
di: Kim, Joongwon, et al.
Pubblicazione: (2024)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
di: Bentham, Oliver, et al.
Pubblicazione: (2026)
di: Bentham, Oliver, et al.
Pubblicazione: (2026)
Mechanistic?
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
di: Chen, Yuefei, et al.
Pubblicazione: (2025)
di: Chen, Yuefei, et al.
Pubblicazione: (2025)
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
di: Miranda, Lester James V., et al.
Pubblicazione: (2024)
di: Miranda, Lester James V., et al.
Pubblicazione: (2024)
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
di: Held, William, et al.
Pubblicazione: (2025)
di: Held, William, et al.
Pubblicazione: (2025)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
Better Aligned with Survey Respondents or Training Data? Unveiling Political Leanings of LLMs on U.S. Supreme Court Cases
di: Xu, Shanshan, et al.
Pubblicazione: (2025)
di: Xu, Shanshan, et al.
Pubblicazione: (2025)
Understanding the Logic of Direct Preference Alignment through Logic
di: Richardson, Kyle, et al.
Pubblicazione: (2024)
di: Richardson, Kyle, et al.
Pubblicazione: (2024)
Promptly Predicting Structures: The Return of Inference
di: Mehta, Maitrey, et al.
Pubblicazione: (2024)
di: Mehta, Maitrey, et al.
Pubblicazione: (2024)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Continual Dialogue State Tracking via Example-Guided Question Answering
di: Cho, Hyundong, et al.
Pubblicazione: (2023)
di: Cho, Hyundong, et al.
Pubblicazione: (2023)
LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
di: Ravisankar, Kartik, et al.
Pubblicazione: (2025)
di: Ravisankar, Kartik, et al.
Pubblicazione: (2025)
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
In-Context Example Ordering Guided by Label Distributions
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Are you going to finish that? A Practical Study of the Partial Token Problem
di: Xu, Hao, et al.
Pubblicazione: (2026)
di: Xu, Hao, et al.
Pubblicazione: (2026)
Distillation versus Contrastive Learning: How to Train Your Rerankers
di: Xu, Zhichao, et al.
Pubblicazione: (2025)
di: Xu, Zhichao, et al.
Pubblicazione: (2025)
State Space Models are Strong Text Rerankers
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
An Empirical Investigation of Matrix Factorization Methods for Pre-trained Transformers
di: Gupta, Ashim, et al.
Pubblicazione: (2024)
di: Gupta, Ashim, et al.
Pubblicazione: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)
di: Hase, Peter, et al.
Pubblicazione: (2024)
Improving Clinical Diagnosis with Counterfactual Multi-Agent Reasoning
di: You, Zhiwen, et al.
Pubblicazione: (2026)
di: You, Zhiwen, et al.
Pubblicazione: (2026)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
di: Palta, Shramay, et al.
Pubblicazione: (2025)
di: Palta, Shramay, et al.
Pubblicazione: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
di: Gupta, Ashim, et al.
Pubblicazione: (2023)
di: Gupta, Ashim, et al.
Pubblicazione: (2023)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion
di: Mehta, Maitrey, et al.
Pubblicazione: (2026)
di: Mehta, Maitrey, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On Linear Representations and Pretraining Data Frequency in Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2025) -
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
di: Merrill, William, et al.
Pubblicazione: (2024) -
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
di: Nadkarni, Rahul, et al.
Pubblicazione: (2025) -
Estimating the Causal Effect of Early ArXiving on Paper Acceptance
di: Elazar, Yanai, et al.
Pubblicazione: (2023) -
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
di: Elazar, Yanai, et al.
Pubblicazione: (2026)