Gespeichert in:
| Hauptverfasser: | Luo, Haoyan, Zarlenga, Mateo Espinosa, Jamnik, Mateja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.06342 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Digging Deeper: Learning Multi-Level Concept Hierarchies
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
Hierarchical Concept-based Interpretable Models
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
Understanding Inter-Concept Relationships in Concept-Based Models
von: Raman, Naveen, et al.
Veröffentlicht: (2024)
von: Raman, Naveen, et al.
Veröffentlicht: (2024)
Foundations of Interpretable Models
von: Barbiero, Pietro, et al.
Veröffentlicht: (2025)
von: Barbiero, Pietro, et al.
Veröffentlicht: (2025)
Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2025)
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2025)
Do Concept Bottleneck Models Respect Localities?
von: Raman, Naveen, et al.
Veröffentlicht: (2024)
von: Raman, Naveen, et al.
Veröffentlicht: (2024)
Actionable Interpretability Must Be Defined in Terms of Symmetries
von: Barbiero, Pietro, et al.
Veröffentlicht: (2026)
von: Barbiero, Pietro, et al.
Veröffentlicht: (2026)
Efficient Bias Mitigation Without Privileged Information
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2024)
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2024)
Learning to Receive Help: Intervention-Aware Concept Embedding Models
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2023)
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2023)
End-to-End Ontology Learning with Large Language Models
von: Lo, Andy, et al.
Veröffentlicht: (2024)
von: Lo, Andy, et al.
Veröffentlicht: (2024)
Don't Start Over: A Cost-Effective Framework for Migrating Personalized Prompts Between LLMs
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
Interpretable Neural-Symbolic Concept Reasoning
von: Barbiero, Pietro, et al.
Veröffentlicht: (2023)
von: Barbiero, Pietro, et al.
Veröffentlicht: (2023)
Don't Pay Attention
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning
von: He, Yongquan, et al.
Veröffentlicht: (2024)
von: He, Yongquan, et al.
Veröffentlicht: (2024)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
von: Hernandez, Adriano
Veröffentlicht: (2024)
von: Hernandez, Adriano
Veröffentlicht: (2024)
Don't Throw Away Your Pretrained Model
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
Think, But Don't Overthink: Reproducing Recursive Language Models
von: Wang, Daren
Veröffentlicht: (2026)
von: Wang, Daren
Veröffentlicht: (2026)
Hatevolution: What Static Benchmarks Don't Tell Us
von: Di Bonaventura, Chiara, et al.
Veröffentlicht: (2025)
von: Di Bonaventura, Chiara, et al.
Veröffentlicht: (2025)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
von: Roy, Debjyoti Saha, et al.
Veröffentlicht: (2024)
von: Roy, Debjyoti Saha, et al.
Veröffentlicht: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
Reasoning Models Reason Well, Until They Don't
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
von: Wang, Yuhang, et al.
Veröffentlicht: (2024)
von: Wang, Yuhang, et al.
Veröffentlicht: (2024)
From Understanding to Utilization: A Survey on Explainability for Large Language Models
von: Luo, Haoyan, et al.
Veröffentlicht: (2024)
von: Luo, Haoyan, et al.
Veröffentlicht: (2024)
Tuning Language Models by Mixture-of-Depths Ensemble
von: Luo, Haoyan, et al.
Veröffentlicht: (2024)
von: Luo, Haoyan, et al.
Veröffentlicht: (2024)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
Language Models Don't Learn the Physical Manifestation of Language
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Don't Walk the Line: Boundary Guidance for Filtered Generation
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Don't Act Blindly: Robust GUI Automation via Action-Effect Verification and Self-Correction
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
Frictional Agent Alignment Framework: Slow Down and Don't Break Things
von: Nath, Abhijnan, et al.
Veröffentlicht: (2025)
von: Nath, Abhijnan, et al.
Veröffentlicht: (2025)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
von: Bafna, Niyati, et al.
Veröffentlicht: (2024)
von: Bafna, Niyati, et al.
Veröffentlicht: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
ROAST: Rollout-based On-distribution Activation Steering Technique
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Digging Deeper: Learning Multi-Level Concept Hierarchies
von: Hill, Oscar, et al.
Veröffentlicht: (2026) -
Hierarchical Concept-based Interpretable Models
von: Hill, Oscar, et al.
Veröffentlicht: (2026) -
Understanding Inter-Concept Relationships in Concept-Based Models
von: Raman, Naveen, et al.
Veröffentlicht: (2024) -
Foundations of Interpretable Models
von: Barbiero, Pietro, et al.
Veröffentlicht: (2025) -
Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
von: Zarlenga, Mateo Espinosa, et al.
Veröffentlicht: (2025)