Language Models Can Predict Their Own Behavior
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ashok, Dhananjay, May, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Little Human Data Goes A Long Way
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Can Interpretation Predict Behavior on Unseen Data?
von: Li, Victoria R., et al.
Veröffentlicht: (2025)
von: Li, Victoria R., et al.
Veröffentlicht: (2025)
Large Language Models for Travel Behavior Prediction
von: Mo, Baichuan, et al.
Veröffentlicht: (2023)
von: Mo, Baichuan, et al.
Veröffentlicht: (2023)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
von: Riachi, Roland, et al.
Veröffentlicht: (2025)
Can VLMs Recall Factual Associations From Visual References?
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
von: Liu, Xiaoze, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2026)
Can Language Models Explain Their Own Classification Behavior?
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
von: Barbulescu, George-Octavian, et al.
Veröffentlicht: (2024)
von: Barbulescu, George-Octavian, et al.
Veröffentlicht: (2024)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
PLDR-LLMs Learn A Generalizable Tensor Operator That Can Replace Its Own Deep Neural Net At Inference
von: Gokden, Burc
Veröffentlicht: (2025)
von: Gokden, Burc
Veröffentlicht: (2025)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
von: Choi, Yunho, et al.
Veröffentlicht: (2026)
von: Choi, Yunho, et al.
Veröffentlicht: (2026)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
von: Gowda, Thamme, et al.
Veröffentlicht: (2021)
von: Gowda, Thamme, et al.
Veröffentlicht: (2021)
AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
von: Beux, Yann Le, et al.
Veröffentlicht: (2025)
von: Beux, Yann Le, et al.
Veröffentlicht: (2025)
Can Language Models Discover Scaling Laws?
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
Can Large Language Models Reason and Plan?
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
von: Kambhampati, Subbarao
Veröffentlicht: (2024)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Optimizing Diversity and Quality through Base-Aligned Model Collaboration
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2021)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2021)
Can Large Language Models Understand Intermediate Representations in Compilers?
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
Can Large Language Models Infer Causation from Correlation?
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
PARAMANU-GANITA: Can Small Math Language Models Rival with Large Language Models on Mathematical Reasoning?
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Correcting Large Language Model Behavior via Influence Function
von: Zhang, Han, et al.
Veröffentlicht: (2024)
von: Zhang, Han, et al.
Veröffentlicht: (2024)
Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
von: King, Jack, et al.
Veröffentlicht: (2026)
von: King, Jack, et al.
Veröffentlicht: (2026)
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Controllable Text Generation in the Instruction-Tuning Era
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Molecules?
von: Sadeghi, Shaghayegh, et al.
Veröffentlicht: (2024)
von: Sadeghi, Shaghayegh, et al.
Veröffentlicht: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
von: Mayne, Harry, et al.
Veröffentlicht: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
von: Jacobi, Jonathan, et al.
Veröffentlicht: (2025)
von: Jacobi, Jonathan, et al.
Veröffentlicht: (2025)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
von: Sastre, Ignacio, et al.
Veröffentlicht: (2025)
von: Sastre, Ignacio, et al.
Veröffentlicht: (2025)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
von: Ye, Jiasheng, et al.
Veröffentlicht: (2023)
von: Ye, Jiasheng, et al.
Veröffentlicht: (2023)
PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
von: Xiao, Yijia, et al.
Veröffentlicht: (2023)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
CPLLM: Clinical Prediction with Large Language Models
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2023)
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2023)
Collaborative Performance Prediction for Large Language Models
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Little Human Data Goes A Long Way
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024) -
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026) -
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025) -
Can Interpretation Predict Behavior on Unseen Data?
von: Li, Victoria R., et al.
Veröffentlicht: (2025) -
Large Language Models for Travel Behavior Prediction
von: Mo, Baichuan, et al.
Veröffentlicht: (2023)