Salvato in:
| Autori principali: | Shin, Changho, Yan, Xinya, Jo, Suenggwan, Cho, Sungjun, Chaudhuri, Shourjo Aditya, Sala, Frederic |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.18693 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
di: Cho, Sungjun, et al.
Pubblicazione: (2025)
di: Cho, Sungjun, et al.
Pubblicazione: (2025)
Weak-to-Strong Generalization Through the Data-Centric Lens
di: Shin, Changho, et al.
Pubblicazione: (2024)
di: Shin, Changho, et al.
Pubblicazione: (2024)
Personalize Your LLM: Fake it then Align it
di: Zhang, Yijing, et al.
Pubblicazione: (2025)
di: Zhang, Yijing, et al.
Pubblicazione: (2025)
Zero-Shot Robustification of Zero-Shot Models
di: Adila, Dyah, et al.
Pubblicazione: (2023)
di: Adila, Dyah, et al.
Pubblicazione: (2023)
Is Free Self-Alignment Possible?
di: Adila, Dyah, et al.
Pubblicazione: (2024)
di: Adila, Dyah, et al.
Pubblicazione: (2024)
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
di: Shin, Changho, et al.
Pubblicazione: (2024)
di: Shin, Changho, et al.
Pubblicazione: (2024)
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
di: Cho, Sungjun, et al.
Pubblicazione: (2025)
di: Cho, Sungjun, et al.
Pubblicazione: (2025)
Multimodal Data Curation via Object Detection and Filter Ensembles
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
di: Huang, Tzu-Heng, et al.
Pubblicazione: (2024)
CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
di: Zhao, Jitian, et al.
Pubblicazione: (2026)
di: Zhao, Jitian, et al.
Pubblicazione: (2026)
Weight Updates as Activation Shifts: A Principled Framework for Steering
di: Adila, Dyah, et al.
Pubblicazione: (2026)
di: Adila, Dyah, et al.
Pubblicazione: (2026)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
di: Karnik, Sathwik, et al.
Pubblicazione: (2025)
di: Karnik, Sathwik, et al.
Pubblicazione: (2025)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
di: Kim, Seyeon, et al.
Pubblicazione: (2024)
di: Kim, Seyeon, et al.
Pubblicazione: (2024)
Federated Time Series Generation on Feature and Temporally Misaligned Data
di: Soi, Zhi Wen, et al.
Pubblicazione: (2024)
di: Soi, Zhi Wen, et al.
Pubblicazione: (2024)
Temporal Misalignment in ANN-SNN Conversion and Its Mitigation via Probabilistic Spiking Neurons
di: Bojković, Velibor, et al.
Pubblicazione: (2025)
di: Bojković, Velibor, et al.
Pubblicazione: (2025)
Test-Time Scaling Makes Overtraining Compute-Optimal
di: Roberts, Nicholas, et al.
Pubblicazione: (2026)
di: Roberts, Nicholas, et al.
Pubblicazione: (2026)
Product Manifold Representations for Learning on Biological Pathways
di: McNeela, Daniel, et al.
Pubblicazione: (2024)
di: McNeela, Daniel, et al.
Pubblicazione: (2024)
Convergent Linear Representations of Emergent Misalignment
di: Soligo, Anna, et al.
Pubblicazione: (2025)
di: Soligo, Anna, et al.
Pubblicazione: (2025)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
di: Cha, Sungmin, et al.
Pubblicazione: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
di: Liang, Kaiqu, et al.
Pubblicazione: (2025)
di: Liang, Kaiqu, et al.
Pubblicazione: (2025)
Steering an Active Learning Workflow Towards Novel Materials Discovery via Queue Prioritization
di: Schwarting, Marcus, et al.
Pubblicazione: (2025)
di: Schwarting, Marcus, et al.
Pubblicazione: (2025)
Imitation Learning from a Single Temporally Misaligned Video
di: Huey, William, et al.
Pubblicazione: (2025)
di: Huey, William, et al.
Pubblicazione: (2025)
Temporal Misalignment Attacks against Multimodal Perception in Autonomous Driving
di: Shahriar, Md Hasan, et al.
Pubblicazione: (2025)
di: Shahriar, Md Hasan, et al.
Pubblicazione: (2025)
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction
di: Dsouza, Amanda, et al.
Pubblicazione: (2024)
di: Dsouza, Amanda, et al.
Pubblicazione: (2024)
Curvature-Guided LoRA: Steering in the pretrained NTK subspace
di: Zheng, Frédéric, et al.
Pubblicazione: (2026)
di: Zheng, Frédéric, et al.
Pubblicazione: (2026)
Mitigating Cognitive Inertia in Large Reasoning Models via Latent Spike Steering
di: Lee, Seojin, et al.
Pubblicazione: (2026)
di: Lee, Seojin, et al.
Pubblicazione: (2026)
MEG-GPT: A transformer-based foundation model for magnetoencephalography data
di: Huang, Rukuang, et al.
Pubblicazione: (2025)
di: Huang, Rukuang, et al.
Pubblicazione: (2025)
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Steering Pretrained Drafters during Speculative Decoding
di: Berdoz, Frédéric, et al.
Pubblicazione: (2025)
di: Berdoz, Frédéric, et al.
Pubblicazione: (2025)
COSMOS: Predictable and Cost-Effective Adaptation of LLMs
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
di: Wang, Jiayu, et al.
Pubblicazione: (2025)
Unveiling the Potential of Superexpressive Networks in Implicit Neural Representations
di: Mudiyanselage, Uvini Balasuriya, et al.
Pubblicazione: (2025)
di: Mudiyanselage, Uvini Balasuriya, et al.
Pubblicazione: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
di: Hahm, Dongyoon, et al.
Pubblicazione: (2025)
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
di: Luan, Yao, et al.
Pubblicazione: (2025)
di: Luan, Yao, et al.
Pubblicazione: (2025)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
di: Han, Sungjun, et al.
Pubblicazione: (2024)
di: Han, Sungjun, et al.
Pubblicazione: (2024)
Tackling Time-Series Forecasting Generalization via Mitigating Concept Drift
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2025)
Expressivity-Efficiency Tradeoffs for Hybrid Sequence Models
di: Cooper, John, et al.
Pubblicazione: (2026)
di: Cooper, John, et al.
Pubblicazione: (2026)
Predicting LLM Reasoning Performance with Small Proxy Model
di: Koh, Woosung, et al.
Pubblicazione: (2025)
di: Koh, Woosung, et al.
Pubblicazione: (2025)
Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval
di: Braun, Guillaume, et al.
Pubblicazione: (2026)
di: Braun, Guillaume, et al.
Pubblicazione: (2026)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
di: Cai, Yichao, et al.
Pubblicazione: (2025)
di: Cai, Yichao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
di: Cho, Sungjun, et al.
Pubblicazione: (2025) -
Weak-to-Strong Generalization Through the Data-Centric Lens
di: Shin, Changho, et al.
Pubblicazione: (2024) -
Personalize Your LLM: Fake it then Align it
di: Zhang, Yijing, et al.
Pubblicazione: (2025) -
Zero-Shot Robustification of Zero-Shot Models
di: Adila, Dyah, et al.
Pubblicazione: (2023) -
Is Free Self-Alignment Possible?
di: Adila, Dyah, et al.
Pubblicazione: (2024)