M-PACE: Mother Child Framework for Multimodal Compliance
Fuente:
arXiv
Guardado en:
| Autores principales: | Verma, Shreyash, Kesari, Amit, Trivedi, Vinayak, Purwar, Anupam, Jamidar, Ratnesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
por: Bandraupalli, Srihari, et al.
Publicado: (2025)
por: Bandraupalli, Srihari, et al.
Publicado: (2025)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
por: Darabi, Nastaran, et al.
Publicado: (2026)
por: Darabi, Nastaran, et al.
Publicado: (2026)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
por: Kumar, Divake, et al.
Publicado: (2026)
por: Kumar, Divake, et al.
Publicado: (2026)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
por: Jin, Yiqiao, et al.
Publicado: (2024)
por: Jin, Yiqiao, et al.
Publicado: (2024)
Dynamic semantic VSLAM with known and unknown objects
por: Gu, Sanghyoup, et al.
Publicado: (2024)
por: Gu, Sanghyoup, et al.
Publicado: (2024)
M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
por: Wang, Hongyu, et al.
Publicado: (2024)
por: Wang, Hongyu, et al.
Publicado: (2024)
See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation
por: Rizwan, Naquee, et al.
Publicado: (2026)
por: Rizwan, Naquee, et al.
Publicado: (2026)
EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
por: Xing, Shangyu, et al.
Publicado: (2024)
por: Xing, Shangyu, et al.
Publicado: (2024)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
por: Wan, Zhongwei, et al.
Publicado: (2024)
por: Wan, Zhongwei, et al.
Publicado: (2024)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
por: Qi, Ji, et al.
Publicado: (2023)
por: Qi, Ji, et al.
Publicado: (2023)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
por: Dalal, Dwip, et al.
Publicado: (2025)
por: Dalal, Dwip, et al.
Publicado: (2025)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
por: Wang, Xiangfeng, et al.
Publicado: (2025)
por: Wang, Xiangfeng, et al.
Publicado: (2025)
Interpretable Multimodal Framework for Human-Centered Street Assessment: Integrating Visual-Language Models for Perceptual Urban Diagnostics
por: Lan, HaoTian
Publicado: (2025)
por: Lan, HaoTian
Publicado: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
por: Wang, Kangsheng, et al.
Publicado: (2025)
por: Wang, Kangsheng, et al.
Publicado: (2025)
Pose-Based Sign Language Appearance Transfer
por: Moryossef, Amit, et al.
Publicado: (2024)
por: Moryossef, Amit, et al.
Publicado: (2024)
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
por: Fu, Rao, et al.
Publicado: (2024)
por: Fu, Rao, et al.
Publicado: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
por: Wu, Te-Lin, et al.
Publicado: (2021)
por: Wu, Te-Lin, et al.
Publicado: (2021)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
por: Gan, Chengguang, et al.
Publicado: (2025)
por: Gan, Chengguang, et al.
Publicado: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
por: Verma, Gaurav, et al.
Publicado: (2024)
por: Verma, Gaurav, et al.
Publicado: (2024)
Ham2Pose: Animating Sign Language Notation into Pose Sequences
por: Shalev-Arkushin, Rotem, et al.
Publicado: (2022)
por: Shalev-Arkushin, Rotem, et al.
Publicado: (2022)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
por: Lim, Qi Zhi, et al.
Publicado: (2025)
por: Lim, Qi Zhi, et al.
Publicado: (2025)
Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation
por: Qin, Zhi, et al.
Publicado: (2025)
por: Qin, Zhi, et al.
Publicado: (2025)
MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
por: Belouadi, Jonas, et al.
Publicado: (2025)
por: Belouadi, Jonas, et al.
Publicado: (2025)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
por: Pantazopoulos, Georgios, et al.
Publicado: (2024)
por: Pantazopoulos, Georgios, et al.
Publicado: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
Advancing Toward Robust and Scalable Fingerprint Orientation Estimation: From Gradients to Deep Learning
por: Trivedi, Amit Kumar, et al.
Publicado: (2020)
por: Trivedi, Amit Kumar, et al.
Publicado: (2020)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2024)
por: Jiang, Chaoya, et al.
Publicado: (2024)
LaRe: Latent Refocusing for Multimodal Reasoning
por: Ma, Jizheng, et al.
Publicado: (2025)
por: Ma, Jizheng, et al.
Publicado: (2025)
Dual-branch Prompting for Multimodal Machine Translation
por: Wang, Jie, et al.
Publicado: (2025)
por: Wang, Jie, et al.
Publicado: (2025)
Open-Vocabulary Federated Learning with Multimodal Prototyping
por: Zeng, Huimin, et al.
Publicado: (2024)
por: Zeng, Huimin, et al.
Publicado: (2024)
Grounding Partially-Defined Events in Multimodal Data
por: Sanders, Kate, et al.
Publicado: (2024)
por: Sanders, Kate, et al.
Publicado: (2024)
Cooperative Sentiment Agents for Multimodal Sentiment Analysis
por: Wang, Shanmin, et al.
Publicado: (2024)
por: Wang, Shanmin, et al.
Publicado: (2024)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
por: Gao, Timin, et al.
Publicado: (2024)
por: Gao, Timin, et al.
Publicado: (2024)
Maya: An Instruction Finetuned Multilingual Multimodal Model
por: Alam, Nahid, et al.
Publicado: (2024)
por: Alam, Nahid, et al.
Publicado: (2024)
Reinforcing Multimodal Reasoning Against Visual Degradation
por: Liu, Rui, et al.
Publicado: (2026)
por: Liu, Rui, et al.
Publicado: (2026)
UEval: A Benchmark for Unified Multimodal Generation
por: Li, Bo, et al.
Publicado: (2026)
por: Li, Bo, et al.
Publicado: (2026)
Veagle: Advancements in Multimodal Representation Learning
por: Chawla, Rajat, et al.
Publicado: (2024)
por: Chawla, Rajat, et al.
Publicado: (2024)
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
por: Khan, Shaharukh, et al.
Publicado: (2025)
por: Khan, Shaharukh, et al.
Publicado: (2025)
Docopilot: Improving Multimodal Models for Document-Level Understanding
por: Duan, Yuchen, et al.
Publicado: (2025)
por: Duan, Yuchen, et al.
Publicado: (2025)
Ejemplares similares
-
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
por: Bandraupalli, Srihari, et al.
Publicado: (2025) -
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
por: Darabi, Nastaran, et al.
Publicado: (2026) -
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
por: Kumar, Divake, et al.
Publicado: (2026) -
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
por: Jin, Yiqiao, et al.
Publicado: (2024) -
Dynamic semantic VSLAM with known and unknown objects
por: Gu, Sanghyoup, et al.
Publicado: (2024)