NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiaxuan, Mo, Junwen, Vo, MinhDuc, Sugimoto, Akihiro, Nakayama, Hideki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
von: Li, Jiaxuan, et al.
Veröffentlicht: (2023)
von: Li, Jiaxuan, et al.
Veröffentlicht: (2023)
Persistent Test-time Adaptation in Recurring Testing Scenarios
von: Hoang, Trung-Hieu, et al.
Veröffentlicht: (2023)
von: Hoang, Trung-Hieu, et al.
Veröffentlicht: (2023)
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
Interactive Masked Image Modeling for Multimodal Object Detection in Remote Sensing
von: Vu, Minh-Duc, et al.
Veröffentlicht: (2024)
von: Vu, Minh-Duc, et al.
Veröffentlicht: (2024)
TESO: Online Tracking of Essential Matrix by Stochastic Optimization
von: Moravec, Jaroslav, et al.
Veröffentlicht: (2026)
von: Moravec, Jaroslav, et al.
Veröffentlicht: (2026)
Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
von: Hoang, Trung-Hieu, et al.
Veröffentlicht: (2023)
von: Hoang, Trung-Hieu, et al.
Veröffentlicht: (2023)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation
von: Neau, Maëlic, et al.
Veröffentlicht: (2025)
von: Neau, Maëlic, et al.
Veröffentlicht: (2025)
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
von: Chen, Junwen, et al.
Veröffentlicht: (2025)
von: Chen, Junwen, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
von: Xiang, Xinhao, et al.
Veröffentlicht: (2025)
von: Xiang, Xinhao, et al.
Veröffentlicht: (2025)
EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing
von: Jiang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Hongxiang, et al.
Veröffentlicht: (2025)
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
von: Tran, Tuyen, et al.
Veröffentlicht: (2025)
InstructAttribute: Fine-grained Object Attributes editing with Instruction
von: Yin, Xingxi, et al.
Veröffentlicht: (2025)
von: Yin, Xingxi, et al.
Veröffentlicht: (2025)
Enhanced Data Transfer Cooperating with Artificial Triplets for Scene Graph Generation
von: Chu, KuanChao, et al.
Veröffentlicht: (2024)
von: Chu, KuanChao, et al.
Veröffentlicht: (2024)
Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025)
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
von: Vo, Khang H. N., et al.
Veröffentlicht: (2025)
von: Vo, Khang H. N., et al.
Veröffentlicht: (2025)
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions
von: Tan, Zhijie, et al.
Veröffentlicht: (2025)
von: Tan, Zhijie, et al.
Veröffentlicht: (2025)
EDGE-Shield: Efficient Denoising-staGE Shield for Violative Content Filtering via Scalable Reference-Based Matching
von: Taniguchi, Takara, et al.
Veröffentlicht: (2026)
von: Taniguchi, Takara, et al.
Veröffentlicht: (2026)
SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
von: Grimal, Paul, et al.
Veröffentlicht: (2025)
von: Grimal, Paul, et al.
Veröffentlicht: (2025)
GoDe: Gaussians on Demand for Progressive Level of Detail and Scalable Compression
von: Di Sario, Francesco, et al.
Veröffentlicht: (2025)
von: Di Sario, Francesco, et al.
Veröffentlicht: (2025)
PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2026)
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2026)
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
von: Thakkar, Parth, et al.
Veröffentlicht: (2025)
von: Thakkar, Parth, et al.
Veröffentlicht: (2025)
Anatomical Attention Alignment representation for Radiology Report Generation
von: Nguyen, Quang Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang Vinh, et al.
Veröffentlicht: (2025)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
von: Masui, Kento, et al.
Veröffentlicht: (2024)
von: Masui, Kento, et al.
Veröffentlicht: (2024)
MureObjectStitch: Multi-reference Image Composition
von: Chen, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxuan, et al.
Veröffentlicht: (2024)
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
von: Xiu, Lixin, et al.
Veröffentlicht: (2026)
von: Xiu, Lixin, et al.
Veröffentlicht: (2026)
REACT: Real-time Efficiency and Accuracy Compromise for Tradeoffs in Scene Graph Generation
von: Neau, Maëlic, et al.
Veröffentlicht: (2024)
von: Neau, Maëlic, et al.
Veröffentlicht: (2024)
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
von: Nguyen, Dung, et al.
Veröffentlicht: (2025)
von: Nguyen, Dung, et al.
Veröffentlicht: (2025)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
von: Wang, Muyao, et al.
Veröffentlicht: (2026)
von: Wang, Muyao, et al.
Veröffentlicht: (2026)
Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
von: Shen, Yutao, et al.
Veröffentlicht: (2025)
von: Shen, Yutao, et al.
Veröffentlicht: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
Material Fingerprinting: Identifying and Predicting Perceptual Attributes of Material Appearance
von: Filip, Jiri, et al.
Veröffentlicht: (2024)
von: Filip, Jiri, et al.
Veröffentlicht: (2024)
Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
von: Yan, Chongshang, et al.
Veröffentlicht: (2025)
von: Yan, Chongshang, et al.
Veröffentlicht: (2025)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2025)
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2025)
ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Minh Anh, et al.
Veröffentlicht: (2026)
Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation
von: Li, Weihang, et al.
Veröffentlicht: (2025)
von: Li, Weihang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
von: Li, Jiaxuan, et al.
Veröffentlicht: (2023) -
Persistent Test-time Adaptation in Recurring Testing Scenarios
von: Hoang, Trung-Hieu, et al.
Veröffentlicht: (2023) -
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
von: Vo, Dinh-Khoi, et al.
Veröffentlicht: (2025) -
Interactive Masked Image Modeling for Multimodal Object Detection in Remote Sensing
von: Vu, Minh-Duc, et al.
Veröffentlicht: (2024) -
TESO: Online Tracking of Essential Matrix by Stochastic Optimization
von: Moravec, Jaroslav, et al.
Veröffentlicht: (2026)