MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Yanrui, Fan, Fenglei, Zhao, Sendong, Cao, Jiawei, Liu, Ting, Qin, Bing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability
von: Du, Yanrui, et al.
Veröffentlicht: (2024)
von: Du, Yanrui, et al.
Veröffentlicht: (2024)
Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
von: Du, Yanrui, et al.
Veröffentlicht: (2024)
von: Du, Yanrui, et al.
Veröffentlicht: (2024)
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
von: Du, Yanrui, et al.
Veröffentlicht: (2025)
von: Du, Yanrui, et al.
Veröffentlicht: (2025)
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text Prompts
von: Guo, Haoqiang, et al.
Veröffentlicht: (2024)
von: Guo, Haoqiang, et al.
Veröffentlicht: (2024)
From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery
von: Chen, Yuhan, et al.
Veröffentlicht: (2023)
von: Chen, Yuhan, et al.
Veröffentlicht: (2023)
AS-ES Learning: Towards Efficient CoT Learning in Small Models
von: Xi, Nuwa, et al.
Veröffentlicht: (2024)
von: Xi, Nuwa, et al.
Veröffentlicht: (2024)
SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation
von: Wang, Haochun, et al.
Veröffentlicht: (2026)
von: Wang, Haochun, et al.
Veröffentlicht: (2026)
MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
von: Aviv, Gilad, et al.
Veröffentlicht: (2025)
von: Aviv, Gilad, et al.
Veröffentlicht: (2025)
Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing
von: Qiang, Zewen, et al.
Veröffentlicht: (2025)
von: Qiang, Zewen, et al.
Veröffentlicht: (2025)
From Latent Signals to Reflection Behavior: Tracing Meta-Cognitive Activation Trajectory in R1-Style LLMs
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
von: Wang, Haochun, et al.
Veröffentlicht: (2023)
von: Wang, Haochun, et al.
Veröffentlicht: (2023)
Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis
von: Wang, Haochun, et al.
Veröffentlicht: (2024)
von: Wang, Haochun, et al.
Veröffentlicht: (2024)
LLMs May Perform MCQA by Selecting the Least Incorrect Option
von: Wang, Haochun, et al.
Veröffentlicht: (2024)
von: Wang, Haochun, et al.
Veröffentlicht: (2024)
MolFusion: Multimodal Fusion Learning for Molecular Representations via Multi-granularity Views
von: Cai, Muzhen, et al.
Veröffentlicht: (2024)
von: Cai, Muzhen, et al.
Veröffentlicht: (2024)
Manifold-based Verbalizer Space Re-embedding for Tuning-free Prompt-based Classification
von: Wang, Haochun, et al.
Veröffentlicht: (2023)
von: Wang, Haochun, et al.
Veröffentlicht: (2023)
S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
von: Du, Yanrui, et al.
Veröffentlicht: (2026)
META-RAG: Meta-Analysis-Inspired Evidence-Re-Ranking Method for Retrieval-Augmented Generation in Evidence-Based Medicine
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection
von: Wang, Haochun, et al.
Veröffentlicht: (2026)
von: Wang, Haochun, et al.
Veröffentlicht: (2026)
PICOs-RAG: PICO-supported Query Rewriting for Retrieval-Augmented Generation in Evidence-Based Medicine
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
ArcAligner: Adaptive Recursive Aligner for Compressed Context Embeddings in RAG
von: Li, Jianbo, et al.
Veröffentlicht: (2026)
von: Li, Jianbo, et al.
Veröffentlicht: (2026)
CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
von: Jiang, Yi, et al.
Veröffentlicht: (2025)
von: Jiang, Yi, et al.
Veröffentlicht: (2025)
OptiSet: Unified Optimizing Set Selection and Ranking for Retrieval-Augmented Generation
von: Jiang, Yi, et al.
Veröffentlicht: (2026)
von: Jiang, Yi, et al.
Veröffentlicht: (2026)
Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models
von: Prucs, Ákos, et al.
Veröffentlicht: (2025)
von: Prucs, Ákos, et al.
Veröffentlicht: (2025)
M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
von: Sun, Mengzhou, et al.
Veröffentlicht: (2025)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
von: Xiong, Kai, et al.
Veröffentlicht: (2023)
von: Xiong, Kai, et al.
Veröffentlicht: (2023)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
Improving the Downstream Performance of Mixture-of-Experts Transformers via Weak Vanilla Transformers
von: Lu, Xin, et al.
Veröffentlicht: (2024)
von: Lu, Xin, et al.
Veröffentlicht: (2024)
Towards Pareto Optimal Throughput in Small Language Model Serving
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
Text Difficulty Study: Do machines behave the same as humans regarding text difficulty?
von: Chen, Bowen, et al.
Veröffentlicht: (2022)
von: Chen, Bowen, et al.
Veröffentlicht: (2022)
Towards Generalizable and Faithful Logic Reasoning over Natural Language via Resolution Refutation
von: Sun, Zhouhao, et al.
Veröffentlicht: (2024)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2024)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
von: Hu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2025)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025)
von: Lu, Xin, et al.
Veröffentlicht: (2025)
Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
Diagnosing and Remedying Knowledge Deficiencies in LLMs via Label-free Curricular Meaningful Learning
von: Xiong, Kai, et al.
Veröffentlicht: (2024)
von: Xiong, Kai, et al.
Veröffentlicht: (2024)
Large Language Models Are Still Misled by Simple Bias Ensembles
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability
von: Du, Yanrui, et al.
Veröffentlicht: (2024) -
Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
von: Du, Yanrui, et al.
Veröffentlicht: (2024) -
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
von: Du, Yanrui, et al.
Veröffentlicht: (2025) -
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
von: Du, Yanrui, et al.
Veröffentlicht: (2023) -
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
von: Du, Yanrui, et al.
Veröffentlicht: (2023)