COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Yizhuo, Chen, Mingkang, Liu, Qiuhua, Weng, Fenghua, Qu, Wanying, Yang, Yue, Jiang, Yugang, Wu, Zuxuan, Fu, Yanwei, Shao, Wenqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
Modeling Spatiotemporal Neural Frames for High Resolution Brain Dynamic
di: Qu, Wanying, et al.
Pubblicazione: (2026)
di: Qu, Wanying, et al.
Pubblicazione: (2026)
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
Revisiting Large Language Model Pruning using Neuron Semantic Attribution
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2024)
di: Weng, Fenghua, et al.
Pubblicazione: (2024)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
Real-Time Multi-Scene Visibility Enhancement for Promoting Navigational Safety of Vessels Under Complex Weather Conditions
di: Liu, Ryan Wen, et al.
Pubblicazione: (2024)
di: Liu, Ryan Wen, et al.
Pubblicazione: (2024)
Trustworthy Distributed AI Systems: Robustness, Privacy, and Governance
di: Wei, Wenqi, et al.
Pubblicazione: (2024)
di: Wei, Wenqi, et al.
Pubblicazione: (2024)
SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
di: Zhang, Lingyun, et al.
Pubblicazione: (2025)
di: Zhang, Lingyun, et al.
Pubblicazione: (2025)
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
di: You, Zuyao, et al.
Pubblicazione: (2026)
di: You, Zuyao, et al.
Pubblicazione: (2026)
Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
di: Hayakawa, Akio, et al.
Pubblicazione: (2025)
di: Hayakawa, Akio, et al.
Pubblicazione: (2025)
GTHNA: Local-global Graph Transformer with Memory Reconstruction for Holistic Node Anomaly Evaluation
di: Li, Mingkang, et al.
Pubblicazione: (2025)
di: Li, Mingkang, et al.
Pubblicazione: (2025)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
di: Li, Albus Yizhuo
Pubblicazione: (2025)
di: Li, Albus Yizhuo
Pubblicazione: (2025)
DCP: Learning Accelerator Dataflow for Neural Network via Propagation
di: Xu, Peng, et al.
Pubblicazione: (2024)
di: Xu, Peng, et al.
Pubblicazione: (2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
di: Wang, Junke, et al.
Pubblicazione: (2024)
di: Wang, Junke, et al.
Pubblicazione: (2024)
COSMO-Bench: A Benchmark for Collaborative SLAM Optimization
di: McGann, Daniel, et al.
Pubblicazione: (2025)
di: McGann, Daniel, et al.
Pubblicazione: (2025)
Safety-aware Causal Representation for Trustworthy Offline Reinforcement Learning in Autonomous Driving
di: Lin, Haohong, et al.
Pubblicazione: (2023)
di: Lin, Haohong, et al.
Pubblicazione: (2023)
Double‐Phase‐Networking Polyimide Hybrid Aerogel with Exceptional Dimensional Stability for Superior Thermal Protection System
di: Chun Liu, et al.
Pubblicazione: (2024)
di: Chun Liu, et al.
Pubblicazione: (2024)
EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
di: Zhang, Zihao, et al.
Pubblicazione: (2025)
di: Zhang, Zihao, et al.
Pubblicazione: (2025)
Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space
di: Liang, Yingping, et al.
Pubblicazione: (2025)
di: Liang, Yingping, et al.
Pubblicazione: (2025)
Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin
di: Hsu, Po-Chun, et al.
Pubblicazione: (2026)
di: Hsu, Po-Chun, et al.
Pubblicazione: (2026)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
di: Xing, Zhen, et al.
Pubblicazione: (2024)
di: Xing, Zhen, et al.
Pubblicazione: (2024)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
di: Weng, Zejia, et al.
Pubblicazione: (2024)
di: Weng, Zejia, et al.
Pubblicazione: (2024)
Computer Vision-Driven Gesture Recognition: Toward Natural and Intuitive Human-Computer
di: Shao, Fenghua, et al.
Pubblicazione: (2024)
di: Shao, Fenghua, et al.
Pubblicazione: (2024)
PromptFusion: Decoupling Stability and Plasticity for Continual Learning
di: Chen, Haoran, et al.
Pubblicazione: (2023)
di: Chen, Haoran, et al.
Pubblicazione: (2023)
Linear-Quadratic Zero-Sum Stochastic Differential Game with Partial Observation
di: Yu, Zhiyong, et al.
Pubblicazione: (2025)
di: Yu, Zhiyong, et al.
Pubblicazione: (2025)
Enhanced Yield Rate of \textsuperscript{229m}Th via Cascade Decay in Storage Rings and Electron Beam Ion Traps
di: Wang, Yumiao, et al.
Pubblicazione: (2026)
di: Wang, Yumiao, et al.
Pubblicazione: (2026)
Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts
di: Zhou, Yunfan, et al.
Pubblicazione: (2025)
di: Zhou, Yunfan, et al.
Pubblicazione: (2025)
COSMO-INR: Complex Sinusoidal Modulation for Implicit Neural Representations
di: Thennakoon, Pandula, et al.
Pubblicazione: (2025)
di: Thennakoon, Pandula, et al.
Pubblicazione: (2025)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
di: Wang, Junke, et al.
Pubblicazione: (2025)
di: Wang, Junke, et al.
Pubblicazione: (2025)
Unsteady Aerodynamics of Bridge Decks in Transient Winds: Vortex Memory Effects and Path‐Dependent Responses
di: Wang Tang, et al.
Pubblicazione: (2025)
di: Wang Tang, et al.
Pubblicazione: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models
di: Tamo, J. Ben, et al.
Pubblicazione: (2026)
di: Tamo, J. Ben, et al.
Pubblicazione: (2026)
Recent Advancements on Spin Engineering Strategies for Highly Efficient Electrocatalytic Oxygen Evolution Reactions
di: Wenli Zhao, et al.
Pubblicazione: (2024)
di: Wenli Zhao, et al.
Pubblicazione: (2024)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
di: Xu, Ziqiang, et al.
Pubblicazione: (2025)
di: Xu, Ziqiang, et al.
Pubblicazione: (2025)
Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
di: Yan, Yanfu, et al.
Pubblicazione: (2025)
di: Yan, Yanfu, et al.
Pubblicazione: (2025)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
di: Chakroborti, Apu Kumar, et al.
Pubblicazione: (2025)
di: Chakroborti, Apu Kumar, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
di: Ding, Yizhuo, et al.
Pubblicazione: (2025) -
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
di: Ding, Yizhuo, et al.
Pubblicazione: (2025) -
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025) -
Modeling Spatiotemporal Neural Frames for High Resolution Brain Dynamic
di: Qu, Wanying, et al.
Pubblicazione: (2026) -
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)