Guardado en:
| Autores principales: | Patil, Vaidehi, Sung, Yi-Lin, Hase, Peter, Peng, Jie, Chen, Tianlong, Bansal, Mohit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2505.01456 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchy-Aware Multimodal Unlearning for Medical AI
por: Wu, Fengli, et al.
Publicado: (2025)
por: Wu, Fengli, et al.
Publicado: (2025)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
por: Patil, Vaidehi, et al.
Publicado: (2025)
por: Patil, Vaidehi, et al.
Publicado: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
por: Sung, Yi-Lin, et al.
Publicado: (2023)
por: Sung, Yi-Lin, et al.
Publicado: (2023)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
por: Lin, Han, et al.
Publicado: (2025)
por: Lin, Han, et al.
Publicado: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
por: Niu, Tianyi, et al.
Publicado: (2025)
por: Niu, Tianyi, et al.
Publicado: (2025)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
por: Yu, Shoubin, et al.
Publicado: (2024)
por: Yu, Shoubin, et al.
Publicado: (2024)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
por: Cheng, Feng, et al.
Publicado: (2024)
por: Cheng, Feng, et al.
Publicado: (2024)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
por: Li, Jialu, et al.
Publicado: (2025)
por: Li, Jialu, et al.
Publicado: (2025)
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
por: Liu, Xuannan, et al.
Publicado: (2024)
por: Liu, Xuannan, et al.
Publicado: (2024)
The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
por: Patil, Vaidehi, et al.
Publicado: (2025)
por: Patil, Vaidehi, et al.
Publicado: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
por: Wang, Han, et al.
Publicado: (2026)
por: Wang, Han, et al.
Publicado: (2026)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
por: Wan, David, et al.
Publicado: (2025)
por: Wan, David, et al.
Publicado: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
por: Kim, Minkyoung, et al.
Publicado: (2024)
por: Kim, Minkyoung, et al.
Publicado: (2024)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
por: Pothiraj, Atin, et al.
Publicado: (2025)
por: Pothiraj, Atin, et al.
Publicado: (2025)
Multimodal Fact-Level Attribution for Verifiable Reasoning
por: Wan, David, et al.
Publicado: (2026)
por: Wan, David, et al.
Publicado: (2026)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
por: Naseh, Ali, et al.
Publicado: (2024)
por: Naseh, Ali, et al.
Publicado: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
por: Wang, Zun, et al.
Publicado: (2024)
por: Wang, Zun, et al.
Publicado: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
por: Zala, Abhay, et al.
Publicado: (2023)
por: Zala, Abhay, et al.
Publicado: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
por: Lin, Han, et al.
Publicado: (2023)
por: Lin, Han, et al.
Publicado: (2023)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
por: Qian, Yusu, et al.
Publicado: (2024)
por: Qian, Yusu, et al.
Publicado: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
por: Lee, Daeun, et al.
Publicado: (2024)
por: Lee, Daeun, et al.
Publicado: (2024)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
por: Ma, David, et al.
Publicado: (2025)
por: Ma, David, et al.
Publicado: (2025)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
por: Rajabi, Navid, et al.
Publicado: (2024)
por: Rajabi, Navid, et al.
Publicado: (2024)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
por: Zong, Yi, et al.
Publicado: (2024)
por: Zong, Yi, et al.
Publicado: (2024)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
por: Guo, Zichun, et al.
Publicado: (2026)
por: Guo, Zichun, et al.
Publicado: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
por: Cheng, Zhili, et al.
Publicado: (2025)
por: Cheng, Zhili, et al.
Publicado: (2025)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
por: Wang, Xiyao, et al.
Publicado: (2024)
por: Wang, Xiyao, et al.
Publicado: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
por: Prasad, Archiki, et al.
Publicado: (2023)
por: Prasad, Archiki, et al.
Publicado: (2023)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
por: Yu, Shoubin, et al.
Publicado: (2026)
por: Yu, Shoubin, et al.
Publicado: (2026)
Ejemplares similares
-
Hierarchy-Aware Multimodal Unlearning for Medical AI
por: Wu, Fengli, et al.
Publicado: (2025) -
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
por: Patil, Vaidehi, et al.
Publicado: (2025) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
por: Sung, Yi-Lin, et al.
Publicado: (2023) -
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
por: Lin, Han, et al.
Publicado: (2025) -
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
por: Niu, Tianyi, et al.
Publicado: (2025)