World Model Robustness via Surprise Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Zollicoffer, Geigh, Chopra, Tanush, Yan, Mingkuan, Ma, Xiaoxu, Eaton, Kenneth, Riedl, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Novelty Detection in Reinforcement Learning with World Models
por: Zollicoffer, Geigh, et al.
Publicado: (2023)
por: Zollicoffer, Geigh, et al.
Publicado: (2023)
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
por: Balloch, Jonathan C., et al.
Publicado: (2024)
por: Balloch, Jonathan C., et al.
Publicado: (2024)
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
por: Eaton, Kenneth, et al.
Publicado: (2024)
por: Eaton, Kenneth, et al.
Publicado: (2024)
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
por: Vu, Minh, et al.
Publicado: (2025)
por: Vu, Minh, et al.
Publicado: (2025)
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
por: Chopra, Tanush, et al.
Publicado: (2024)
por: Chopra, Tanush, et al.
Publicado: (2024)
Topological Signatures of Adversaries in Multimodal Alignments
por: Vu, Minh, et al.
Publicado: (2025)
por: Vu, Minh, et al.
Publicado: (2025)
LoRID: Low-Rank Iterative Diffusion for Adversarial Purification
por: Zollicoffer, Geigh, et al.
Publicado: (2024)
por: Zollicoffer, Geigh, et al.
Publicado: (2024)
LaFA: Latent Feature Attacks on Non-negative Matrix Factorization
por: Vu, Minh, et al.
Publicado: (2024)
por: Vu, Minh, et al.
Publicado: (2024)
MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
por: Zollicoffer, Geigh, et al.
Publicado: (2025)
por: Zollicoffer, Geigh, et al.
Publicado: (2025)
Accelerating Matrix Diagonalization through Decision Transformers with Epsilon-Greedy Optimization
por: Bhatta, Kshitij, et al.
Publicado: (2024)
por: Bhatta, Kshitij, et al.
Publicado: (2024)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
por: Zollicoffer, Geigh, et al.
Publicado: (2024)
por: Zollicoffer, Geigh, et al.
Publicado: (2024)
Hybrid Neural World Models
por: Lakshmanan, Pranav, et al.
Publicado: (2026)
por: Lakshmanan, Pranav, et al.
Publicado: (2026)
Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning
por: Banerjee, Amartya, et al.
Publicado: (2023)
por: Banerjee, Amartya, et al.
Publicado: (2023)
The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
por: Ma, Jiajun, et al.
Publicado: (2024)
por: Ma, Jiajun, et al.
Publicado: (2024)
External Model Motivated Agents: Reinforcement Learning for Enhanced Environment Sampling
por: Bhagat, Rishav, et al.
Publicado: (2024)
por: Bhagat, Rishav, et al.
Publicado: (2024)
Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
por: Balogh, András, et al.
Publicado: (2026)
por: Balogh, András, et al.
Publicado: (2026)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
por: Sharma, Aman, et al.
Publicado: (2026)
por: Sharma, Aman, et al.
Publicado: (2026)
Mesa-Extrapolation: A Weave Position Encoding Method for Enhanced Extrapolation in LLMs
por: Ma, Xin, et al.
Publicado: (2024)
por: Ma, Xin, et al.
Publicado: (2024)
Sanity Checks for Long-Form Hallucination Detection
por: Zollicoffer, Geigh, et al.
Publicado: (2026)
por: Zollicoffer, Geigh, et al.
Publicado: (2026)
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models
por: Farhat, Sean, et al.
Publicado: (2024)
por: Farhat, Sean, et al.
Publicado: (2024)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
por: Hu, Wentao, et al.
Publicado: (2025)
por: Hu, Wentao, et al.
Publicado: (2025)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
por: Khanna, Sarthak, et al.
Publicado: (2025)
por: Khanna, Sarthak, et al.
Publicado: (2025)
Model Agreement via Anchoring
por: Eaton, Eric, et al.
Publicado: (2026)
por: Eaton, Eric, et al.
Publicado: (2026)
On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
por: Joo, Taejong, et al.
Publicado: (2026)
por: Joo, Taejong, et al.
Publicado: (2026)
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
por: Chandhok, Shivam, et al.
Publicado: (2025)
por: Chandhok, Shivam, et al.
Publicado: (2025)
Influence functions and regularity tangents for efficient active learning
por: Eaton, Frederik
Publicado: (2024)
por: Eaton, Frederik
Publicado: (2024)
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
por: Sharma, Aman, et al.
Publicado: (2025)
por: Sharma, Aman, et al.
Publicado: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
por: Sharma, Aman, et al.
Publicado: (2025)
por: Sharma, Aman, et al.
Publicado: (2025)
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
por: Trehan, Dhruv, et al.
Publicado: (2026)
por: Trehan, Dhruv, et al.
Publicado: (2026)
Discovering Reinforcement Learning Interfaces with Large Language Models
por: Jaswal, Akshat Singh, et al.
Publicado: (2026)
por: Jaswal, Akshat Singh, et al.
Publicado: (2026)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
por: Hugessen, Adriana, et al.
Publicado: (2024)
por: Hugessen, Adriana, et al.
Publicado: (2024)
Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction
por: GX-Chen, Anthony, et al.
Publicado: (2024)
por: GX-Chen, Anthony, et al.
Publicado: (2024)
Learning Latent Dynamic Robust Representations for World Models
por: Sun, Ruixiang, et al.
Publicado: (2024)
por: Sun, Ruixiang, et al.
Publicado: (2024)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
por: Liu, Huihan, et al.
Publicado: (2026)
por: Liu, Huihan, et al.
Publicado: (2026)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
por: Agarwal, Dhruv, et al.
Publicado: (2025)
por: Agarwal, Dhruv, et al.
Publicado: (2025)
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
por: Haas, Moritz, et al.
Publicado: (2025)
por: Haas, Moritz, et al.
Publicado: (2025)
Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection
por: Özer, Kadir-Kaan, et al.
Publicado: (2026)
por: Özer, Kadir-Kaan, et al.
Publicado: (2026)
Discrete World Models via Regularization
por: Bizzaro, Davide, et al.
Publicado: (2026)
por: Bizzaro, Davide, et al.
Publicado: (2026)
Building Interpretable Models for Moral Decision-Making
por: Goel, Mayank, et al.
Publicado: (2026)
por: Goel, Mayank, et al.
Publicado: (2026)
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
por: Cheng, Jie, et al.
Publicado: (2024)
por: Cheng, Jie, et al.
Publicado: (2024)
Ejemplares similares
-
Novelty Detection in Reinforcement Learning with World Models
por: Zollicoffer, Geigh, et al.
Publicado: (2023) -
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
por: Balloch, Jonathan C., et al.
Publicado: (2024) -
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
por: Eaton, Kenneth, et al.
Publicado: (2024) -
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
por: Vu, Minh, et al.
Publicado: (2025) -
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
por: Chopra, Tanush, et al.
Publicado: (2024)