Saved in:
| Main Authors: | Jang, Donggon, Cho, Yucheol, Lee, Suin, Kim, Taehyeon, Kim, Dae-Shik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.13881 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maximizing Discrimination Capability of Knowledge Distillation with Energy Function
by: Kim, Seonghak, et al.
Published: (2023)
by: Kim, Seonghak, et al.
Published: (2023)
TexTailor: Customized Text-aligned Texturing via Effective Resampling
by: Lee, Suin, et al.
Published: (2025)
by: Lee, Suin, et al.
Published: (2025)
Robustness-Reinforced Knowledge Distillation with Correlation Distance and Network Pruning
by: Kim, Seonghak, et al.
Published: (2023)
by: Kim, Seonghak, et al.
Published: (2023)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
by: Jang, Won Shik, et al.
Published: (2026)
by: Jang, Won Shik, et al.
Published: (2026)
Semantic Layering in Room Segmentation via LLMs
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Discontinuity-preserving Normal Integration with Auxiliary Edges
by: Kim, Hyomin, et al.
Published: (2024)
by: Kim, Hyomin, et al.
Published: (2024)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
by: Eum, Taehyeon, et al.
Published: (2025)
by: Eum, Taehyeon, et al.
Published: (2025)
Decoding fMRI Data into Captions using Prefix Language Modeling
by: Shen, Vyacheslav, et al.
Published: (2025)
by: Shen, Vyacheslav, et al.
Published: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026)
by: Lee, Youngwan, et al.
Published: (2026)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
by: Kim, Soo Yong, et al.
Published: (2025)
by: Kim, Soo Yong, et al.
Published: (2025)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
Large-scale EM Benchmark for Multi-Organelle Instance Segmentation in the Wild
by: Lu, Yanrui, et al.
Published: (2026)
by: Lu, Yanrui, et al.
Published: (2026)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
by: Yu, Youngjoon, et al.
Published: (2024)
by: Yu, Youngjoon, et al.
Published: (2024)
Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
by: Lee, Seungryong, et al.
Published: (2025)
by: Lee, Seungryong, et al.
Published: (2025)
CEC-MMR: Cross-Entropy Clustering Approach to Multi-Modal Regression
by: Byrski, Krzysztof, et al.
Published: (2025)
by: Byrski, Krzysztof, et al.
Published: (2025)
CXR-LT 2026 Challenge: Projection-Aware Multi-Label and Zero-Shot Chest X-Ray Classification
by: Cho, Juno, et al.
Published: (2026)
by: Cho, Juno, et al.
Published: (2026)
MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
by: Li, Jiachun, et al.
Published: (2026)
by: Li, Jiachun, et al.
Published: (2026)
Treating Motion as Option with Output Selection for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2023)
by: Cho, Suhwan, et al.
Published: (2023)
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
by: Jang, Oh-Tae, et al.
Published: (2025)
by: Jang, Oh-Tae, et al.
Published: (2025)
MMR: Evaluating Reading Ability of Large Multimodal Models
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
PlugTrack: Multi-Perceptive Motion Analysis for Adaptive Fusion in Multi-Object Tracking
by: Kim, Seungjae, et al.
Published: (2025)
by: Kim, Seungjae, et al.
Published: (2025)
BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation
by: AlMughrabi, Ahmad, et al.
Published: (2026)
by: AlMughrabi, Ahmad, et al.
Published: (2026)
PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset
by: Teodoro, Samuel, et al.
Published: (2024)
by: Teodoro, Samuel, et al.
Published: (2024)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
by: Pak, Byeonghyun, et al.
Published: (2024)
by: Pak, Byeonghyun, et al.
Published: (2024)
ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasets
by: Kim, Hoyoung, et al.
Published: (2026)
by: Kim, Hoyoung, et al.
Published: (2026)
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
by: Song, Jeonghyo, et al.
Published: (2025)
by: Song, Jeonghyo, et al.
Published: (2025)
Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction
by: Woo, Sungmin, et al.
Published: (2024)
by: Woo, Sungmin, et al.
Published: (2024)
Frequency-enhanced Multi-granularity Context Network for Efficient Vertebrae Segmentation
by: Shi, Jian, et al.
Published: (2025)
by: Shi, Jian, et al.
Published: (2025)
MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
by: Roy, Aniket, et al.
Published: (2025)
by: Roy, Aniket, et al.
Published: (2025)
Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection
by: Chang, Gyusam, et al.
Published: (2024)
by: Chang, Gyusam, et al.
Published: (2024)
Continual-MEGA: A Large-scale Benchmark for Generalizable Continual Anomaly Detection
by: Lee, Geonu, et al.
Published: (2025)
by: Lee, Geonu, et al.
Published: (2025)
Navigating Data Heterogeneity in Federated Learning A Semi-Supervised Federated Object Detection
by: Kim, Taehyeon, et al.
Published: (2023)
by: Kim, Taehyeon, et al.
Published: (2023)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
by: Kim, SuYeon, et al.
Published: (2026)
by: Kim, SuYeon, et al.
Published: (2026)
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
by: Leng, Sicong, et al.
Published: (2025)
by: Leng, Sicong, et al.
Published: (2025)
Similar Items
-
Maximizing Discrimination Capability of Knowledge Distillation with Energy Function
by: Kim, Seonghak, et al.
Published: (2023) -
TexTailor: Customized Text-aligned Texturing via Effective Resampling
by: Lee, Suin, et al.
Published: (2025) -
Robustness-Reinforced Knowledge Distillation with Correlation Distance and Network Pruning
by: Kim, Seonghak, et al.
Published: (2023) -
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024) -
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
by: Jang, Won Shik, et al.
Published: (2026)