Weird Generalization is Weirdly Brittle
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wanner, Miriam, Collison, Hannah, Jurayj, William, Van Durme, Benjamin, Dredze, Mark, Walden, William |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
A Closer Look at Claim Decomposition
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
von: Wanner, Miriam, et al.
Veröffentlicht: (2024)
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
von: Jurayj, William, et al.
Veröffentlicht: (2025)
von: Jurayj, William, et al.
Veröffentlicht: (2025)
Reasoning Models Will Sometimes Lie About Their Reasoning
von: Walden, William, et al.
Veröffentlicht: (2026)
von: Walden, William, et al.
Veröffentlicht: (2026)
Language Models and Logic Programs for Trustworthy Tax Reasoning
von: Jurayj, William, et al.
Veröffentlicht: (2025)
von: Jurayj, William, et al.
Veröffentlicht: (2025)
One Weird Trick to Untie Landin's Knot
von: Koronkevich, Paulette, et al.
Veröffentlicht: (2025)
von: Koronkevich, Paulette, et al.
Veröffentlicht: (2025)
How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
von: Walden, William, et al.
Veröffentlicht: (2025)
von: Walden, William, et al.
Veröffentlicht: (2025)
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Abe Bohan, et al.
Veröffentlicht: (2024)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
von: Wang, Liaoyaqi, et al.
Veröffentlicht: (2026)
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
von: Mortensen, David R., et al.
Veröffentlicht: (2024)
von: Mortensen, David R., et al.
Veröffentlicht: (2024)
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
von: Wanner, Miriam, et al.
Veröffentlicht: (2025)
von: Wanner, Miriam, et al.
Veröffentlicht: (2025)
DeonticBench: A Benchmark for Reasoning over Rules
von: Dou, Guangyao, et al.
Veröffentlicht: (2026)
von: Dou, Guangyao, et al.
Veröffentlicht: (2026)
Weirding Civilization
von: Herva, Vesa-Pekka, et al.
Veröffentlicht: (2026)
von: Herva, Vesa-Pekka, et al.
Veröffentlicht: (2026)
Weirding Landscapes
von: Herva, Vesa-Pekka, et al.
Veröffentlicht: (2025)
von: Herva, Vesa-Pekka, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
WikiVideo: Article Generation from Multiple Videos
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
Chapter Weird Women
von: Kartsaki, Eirini
Veröffentlicht: (2024)
von: Kartsaki, Eirini
Veröffentlicht: (2024)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
We are a Weird Mob
von: Dredge, Ken
Veröffentlicht: (1998)
von: Dredge, Ken
Veröffentlicht: (1998)
Weird $\mathbb R$-Factorizable Groups
von: Reznichenko, Evgenii, et al.
Veröffentlicht: (2025)
von: Reznichenko, Evgenii, et al.
Veröffentlicht: (2025)
Core: Robust Factual Precision with Informative Sub-Claim Identification
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
von: An, Bang, et al.
Veröffentlicht: (2025)
von: An, Bang, et al.
Veröffentlicht: (2025)
Decomposing Generalization: Models of Generic, Habitual, and Episodic Statements
von: Govindarajan, Venkata Subrahmanyan, et al.
Veröffentlicht: (2019)
von: Govindarajan, Venkata Subrahmanyan, et al.
Veröffentlicht: (2019)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Amuro and Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models
von: Sun, Kaiser, et al.
Veröffentlicht: (2024)
von: Sun, Kaiser, et al.
Veröffentlicht: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
SocialNLI: A Dialogue-Centric Social Inference Dataset
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
von: Deo, Akhil, et al.
Veröffentlicht: (2025)
NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
von: Weir, Nathaniel, et al.
Veröffentlicht: (2022)
von: Weir, Nathaniel, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
von: Wanner, Miriam, et al.
Veröffentlicht: (2024) -
A Closer Look at Claim Decomposition
von: Wanner, Miriam, et al.
Veröffentlicht: (2024) -
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
von: Jurayj, William, et al.
Veröffentlicht: (2025) -
Reasoning Models Will Sometimes Lie About Their Reasoning
von: Walden, William, et al.
Veröffentlicht: (2026) -
Language Models and Logic Programs for Trustworthy Tax Reasoning
von: Jurayj, William, et al.
Veröffentlicht: (2025)