Hierarchy-Aware Multimodal Unlearning for Medical AI
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Fengli, Patil, Vaidehi, Yoon, Jaehong, Zhang, Yue, Bansal, Mohit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
por: Yu, Shoubin, et al.
Publicado: (2024)
por: Yu, Shoubin, et al.
Publicado: (2024)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
por: Patil, Vaidehi, et al.
Publicado: (2025)
por: Patil, Vaidehi, et al.
Publicado: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
por: Lee, Daeun, et al.
Publicado: (2024)
por: Lee, Daeun, et al.
Publicado: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
por: Li, Jialu, et al.
Publicado: (2025)
por: Li, Jialu, et al.
Publicado: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
por: Sung, Yi-Lin, et al.
Publicado: (2023)
por: Sung, Yi-Lin, et al.
Publicado: (2023)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
por: Huang, Yidong, et al.
Publicado: (2025)
por: Huang, Yidong, et al.
Publicado: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
por: Wang, Zun, et al.
Publicado: (2024)
por: Wang, Zun, et al.
Publicado: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
por: Yu, Shoubin, et al.
Publicado: (2026)
por: Yu, Shoubin, et al.
Publicado: (2026)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2026)
por: Wang, Ziyang, et al.
Publicado: (2026)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
por: Wang, Ziyang, et al.
Publicado: (2024)
por: Wang, Ziyang, et al.
Publicado: (2024)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
por: Wang, Zun, et al.
Publicado: (2026)
por: Wang, Zun, et al.
Publicado: (2026)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
por: Wang, Zun, et al.
Publicado: (2025)
por: Wang, Zun, et al.
Publicado: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
por: Niu, Tianyi, et al.
Publicado: (2025)
por: Niu, Tianyi, et al.
Publicado: (2025)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
por: Patil, Vaidehi, et al.
Publicado: (2025)
por: Patil, Vaidehi, et al.
Publicado: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
por: Lin, Han, et al.
Publicado: (2025)
por: Lin, Han, et al.
Publicado: (2025)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
por: Wang, Xiyao, et al.
Publicado: (2024)
por: Wang, Xiyao, et al.
Publicado: (2024)
Multimodal Fact-Level Attribution for Verifiable Reasoning
por: Wan, David, et al.
Publicado: (2026)
por: Wan, David, et al.
Publicado: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
por: Pothiraj, Atin, et al.
Publicado: (2025)
por: Pothiraj, Atin, et al.
Publicado: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
por: Huang, Yidong, et al.
Publicado: (2026)
por: Huang, Yidong, et al.
Publicado: (2026)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
por: Yeo, Woongyeong, et al.
Publicado: (2025)
por: Yeo, Woongyeong, et al.
Publicado: (2025)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
por: Cho, Jaemin, et al.
Publicado: (2024)
por: Cho, Jaemin, et al.
Publicado: (2024)
See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
por: Ananthram, Amith, et al.
Publicado: (2024)
por: Ananthram, Amith, et al.
Publicado: (2024)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
por: Fang, Xianya, et al.
Publicado: (2026)
por: Fang, Xianya, et al.
Publicado: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
por: Zhang, Xiaohui, et al.
Publicado: (2023)
por: Zhang, Xiaohui, et al.
Publicado: (2023)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
por: Prasad, Archiki, et al.
Publicado: (2023)
por: Prasad, Archiki, et al.
Publicado: (2023)
MultiDelete for Multimodal Machine Unlearning
por: Cheng, Jiali, et al.
Publicado: (2023)
por: Cheng, Jiali, et al.
Publicado: (2023)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
por: Zala, Abhay, et al.
Publicado: (2023)
por: Zala, Abhay, et al.
Publicado: (2023)
Ejemplares similares
-
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
por: Yu, Shoubin, et al.
Publicado: (2024) -
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
por: Patil, Vaidehi, et al.
Publicado: (2025) -
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
por: Yu, Shoubin, et al.
Publicado: (2025) -
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
por: Yoon, Jaehong, et al.
Publicado: (2024) -
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
por: Yoon, Jaehong, et al.
Publicado: (2024)