From Data to Behavior: Predicting Unintended Model Behaviors Before Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Mengru, Xu, Zhenqian, Fang, Junfeng, Yao, Yunzhi, Deng, Shumin, Chen, Huajun, Zhang, Ningyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowledge Circuits in Pretrained Transformers
von: Yao, Yunzhi, et al.
Veröffentlicht: (2024)
von: Yao, Yunzhi, et al.
Veröffentlicht: (2024)
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners
von: Yao, Yunzhi, et al.
Veröffentlicht: (2025)
von: Yao, Yunzhi, et al.
Veröffentlicht: (2025)
Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
von: Xu, Ziwen, et al.
Veröffentlicht: (2026)
Editing Conceptual Knowledge for Large Language Models
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs
von: Fang, Jizhan, et al.
Veröffentlicht: (2024)
von: Fang, Jizhan, et al.
Veröffentlicht: (2024)
StructMem: Structured Memory for Long-Horizon Behavior in LLMs
von: Xu, Buqiang, et al.
Veröffentlicht: (2026)
von: Xu, Buqiang, et al.
Veröffentlicht: (2026)
LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities
von: Zhu, Yuqi, et al.
Veröffentlicht: (2023)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2023)
WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Towards A Unified View of Answer Calibration for Multi-Step Reasoning
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
Improving Stance Detection by Leveraging Measurement Knowledge from Social Sciences: A Case Study of Dutch Political Tweets and Traditional Gender Role Division
von: Fang, Qixiang, et al.
Veröffentlicht: (2022)
von: Fang, Qixiang, et al.
Veröffentlicht: (2022)
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
von: Wang, Peng, et al.
Veröffentlicht: (2023)
von: Wang, Peng, et al.
Veröffentlicht: (2023)
OneEdit: A Neural-Symbolic Collaboratively Knowledge Editing System
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
Information Extraction in Low-Resource Scenarios: Survey and Perspective
von: Deng, Shumin, et al.
Veröffentlicht: (2022)
von: Deng, Shumin, et al.
Veröffentlicht: (2022)
ReCode: Updating Code API Knowledge with Reinforcement Learning
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
Predicting Movie Hits Before They Happen with LLMs
von: Agah, Shaghayegh, et al.
Veröffentlicht: (2025)
von: Agah, Shaghayegh, et al.
Veröffentlicht: (2025)
Unmasking the Reality of PII Masking Models: Performance Gaps and the Call for Accountability
von: Singh, Devansh, et al.
Veröffentlicht: (2025)
von: Singh, Devansh, et al.
Veröffentlicht: (2025)
Suicide Phenotyping from Clinical Notes in Safety-Net Psychiatric Hospital Using Multi-Label Classification with Pre-Trained Language Models
von: Li, Zehan, et al.
Veröffentlicht: (2024)
von: Li, Zehan, et al.
Veröffentlicht: (2024)
Signal in the Noise: Decoding the Reality of Airline Service Quality with Large Language Models
von: Dawoud, Ahmed, et al.
Veröffentlicht: (2026)
von: Dawoud, Ahmed, et al.
Veröffentlicht: (2026)
Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
von: Yang, Kai-Cheng, et al.
Veröffentlicht: (2023)
von: Yang, Kai-Cheng, et al.
Veröffentlicht: (2023)
Automating Steering for Safe Multimodal Large Language Models
von: Wu, Lyucheng, et al.
Veröffentlicht: (2025)
von: Wu, Lyucheng, et al.
Veröffentlicht: (2025)
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
von: DeVerna, Matthew R., et al.
Veröffentlicht: (2025)
von: DeVerna, Matthew R., et al.
Veröffentlicht: (2025)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
von: Wang, Mengru, et al.
Veröffentlicht: (2025)
Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
News Source Citing Patterns in AI Search Systems
von: Yang, Kai-Cheng
Veröffentlicht: (2025)
von: Yang, Kai-Cheng
Veröffentlicht: (2025)
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
von: Adak, Sayantan, et al.
Veröffentlicht: (2025)
von: Adak, Sayantan, et al.
Veröffentlicht: (2025)
Corpus-Based Approaches to Igbo Diacritic Restoration
von: Ezeani, Ignatius
Veröffentlicht: (2026)
von: Ezeani, Ignatius
Veröffentlicht: (2026)
Interpretation Gaps in LLM-Assisted Comprehension of Privacy Documents
von: Dewri, Rinku
Veröffentlicht: (2025)
von: Dewri, Rinku
Veröffentlicht: (2025)
GeoOutageKG: A Multimodal Geospatiotemporal Knowledge Graph for Multiresolution Power Outage Analysis
von: Frakes, Ethan, et al.
Veröffentlicht: (2025)
von: Frakes, Ethan, et al.
Veröffentlicht: (2025)
NeuroLit Navigator: A Neurosymbolic Approach to Scholarly Article Searches for Systematic Reviews
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News Recommendation
von: Hu, Beizhe, et al.
Veröffentlicht: (2025)
von: Hu, Beizhe, et al.
Veröffentlicht: (2025)
Ancient Wisdom, Modern Tools: Exploring Retrieval-Augmented LLMs for Ancient Indian Philosophy
von: Mandikal, Priyanka
Veröffentlicht: (2024)
von: Mandikal, Priyanka
Veröffentlicht: (2024)
SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
von: Qiao, Shuofei, et al.
Veröffentlicht: (2026)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2026)
Harmonizing Large Language Models with Collaborative Behavioral Signals for Conversational Recommendation
von: Li, Guanrong, et al.
Veröffentlicht: (2025)
von: Li, Guanrong, et al.
Veröffentlicht: (2025)
Unmasking Superspreaders: Data-Driven Approaches for Identifying and Comparing Key Influencers of Conspiracy Theories on X.com
von: Kramer, Florian, et al.
Veröffentlicht: (2026)
von: Kramer, Florian, et al.
Veröffentlicht: (2026)
Making Language Models Better Tool Learners with Execution Feedback
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
von: Mishra, Shubham, et al.
Veröffentlicht: (2025)
von: Mishra, Shubham, et al.
Veröffentlicht: (2025)
Revisit and Outstrip Entity Alignment: A Perspective of Generative Models
von: Guo, Lingbing, et al.
Veröffentlicht: (2023)
von: Guo, Lingbing, et al.
Veröffentlicht: (2023)
Exploring Model Kinship for Merging Large Language Models
von: Hu, Yedi, et al.
Veröffentlicht: (2024)
von: Hu, Yedi, et al.
Veröffentlicht: (2024)
RaFe: Ranking Feedback Improves Query Rewriting for RAG
von: Mao, Shengyu, et al.
Veröffentlicht: (2024)
von: Mao, Shengyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Knowledge Circuits in Pretrained Transformers
von: Yao, Yunzhi, et al.
Veröffentlicht: (2024) -
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
von: Wang, Mengru, et al.
Veröffentlicht: (2025) -
CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners
von: Yao, Yunzhi, et al.
Veröffentlicht: (2025) -
Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics
von: Xu, Ziwen, et al.
Veröffentlicht: (2026) -
Editing Conceptual Knowledge for Large Language Models
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)