FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhuoyun, Wang, Boxuan, Hu, Jinwei, Huang, Xiaowei, Dong, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models
by: Li, Zhuoyun, et al.
Published: (2026)
by: Li, Zhuoyun, et al.
Published: (2026)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024)
by: Arakelyan, Erik, et al.
Published: (2024)
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
by: Martra, Pere
Published: (2025)
by: Martra, Pere
Published: (2025)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts
by: Wang, Boxuan, et al.
Published: (2026)
by: Wang, Boxuan, et al.
Published: (2026)
Trust-Oriented Adaptive Guardrails for Large Language Models
by: Hu, Jinwei, et al.
Published: (2024)
by: Hu, Jinwei, et al.
Published: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
On The Fragility of Benchmark Contamination Detection in Reasoning Models
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Rethinking Multi-Agent Intelligence Through the Lens of Small-World Networks
by: Wang, Boxuan, et al.
Published: (2025)
by: Wang, Boxuan, et al.
Published: (2025)
Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
by: Wang, Boxuan, et al.
Published: (2025)
by: Wang, Boxuan, et al.
Published: (2025)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
by: Zheng, Guantian
Published: (2026)
by: Zheng, Guantian
Published: (2026)
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
by: Liu, Zhining, et al.
Published: (2026)
by: Liu, Zhining, et al.
Published: (2026)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs
by: Vijayaraj, Raj Krishnan
Published: (2025)
by: Vijayaraj, Raj Krishnan
Published: (2025)
Tool Learning with Foundation Models
by: Qin, Yujia, et al.
Published: (2023)
by: Qin, Yujia, et al.
Published: (2023)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
Weaver: Foundation Models for Creative Writing
by: Wang, Tiannan, et al.
Published: (2024)
by: Wang, Tiannan, et al.
Published: (2024)
Clarify: Improving Model Robustness With Natural Language Corrections
by: Lee, Yoonho, et al.
Published: (2024)
by: Lee, Yoonho, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning for Foundation Models
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
by: Rahman, Roussel, et al.
Published: (2025)
by: Rahman, Roussel, et al.
Published: (2025)
Hierarchical Testing with Rabbit Optimization for Industrial Cyber-Physical Systems
by: Hu, Jinwei, et al.
Published: (2025)
by: Hu, Jinwei, et al.
Published: (2025)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Stabilising Explainability Fragility in Cybersecurity AI: The Impact and Mitigation of Multicollinearity in Public Benchmark Datasets
by: Vourganas, Ioannis J., et al.
Published: (2026)
by: Vourganas, Ioannis J., et al.
Published: (2026)
On the Fragility of Data Attribution When Learning Is Distributed
by: Gao, Xian, et al.
Published: (2026)
by: Gao, Xian, et al.
Published: (2026)
Automated Capability Discovery via Foundation Model Self-Exploration
by: Lu, Cong, et al.
Published: (2025)
by: Lu, Cong, et al.
Published: (2025)
Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models
by: Lu, Cong, et al.
Published: (2024)
by: Lu, Cong, et al.
Published: (2024)
Towards Trustable Language Models: Investigating Information Quality of Large Language Models
by: Rejeleene, Rick, et al.
Published: (2024)
by: Rejeleene, Rick, et al.
Published: (2024)
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
by: Zhang, Pingyue, et al.
Published: (2026)
by: Zhang, Pingyue, et al.
Published: (2026)
Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training
by: Wang, Haoran, et al.
Published: (2026)
by: Wang, Haoran, et al.
Published: (2026)
Conditional Adversarial Fragility in Financial Machine Learning under Macroeconomic Stress
by: Baviskar, Samruddhi
Published: (2025)
by: Baviskar, Samruddhi
Published: (2025)
On the Fragility of AI-Based Channel Decoders under Small Channel Perturbations
by: Lei, Haoyu, et al.
Published: (2026)
by: Lei, Haoyu, et al.
Published: (2026)
How Fragile is Relation Extraction under Entity Replacements?
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
Apple Intelligence Foundation Language Models
by: Gunter, Tom, et al.
Published: (2024)
by: Gunter, Tom, et al.
Published: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Variational Language Concepts for Interpreting Foundation Language Models
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
Similar Items
-
Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models
by: Li, Zhuoyun, et al.
Published: (2026) -
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
by: Arakelyan, Erik, et al.
Published: (2024) -
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
by: Martra, Pere
Published: (2025) -
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026) -
Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts
by: Wang, Boxuan, et al.
Published: (2026)