Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Yejin, Kim, Saejin, Min, Dongjun, Yu, Younjae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024)
by: van Delft, Bastien, et al.
Published: (2024)
An Ordinary Differential Equation Sampler with Stochastic Start for Diffusion Bridge Models
by: Wang, Yuang, et al.
Published: (2024)
by: Wang, Yuang, et al.
Published: (2024)
NODER: Image Sequence Regression Based on Neural Ordinary Differential Equations
by: Bai, Hao, et al.
Published: (2024)
by: Bai, Hao, et al.
Published: (2024)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
A Lightweight U-like Network Utilizing Neural Memory Ordinary Differential Equations for Slimming the Decoder
by: He, Quansong, et al.
Published: (2024)
by: He, Quansong, et al.
Published: (2024)
Teaching Metric Distance to Discrete Autoregressive Language Models
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
by: Ko, Donggeun, et al.
Published: (2024)
by: Ko, Donggeun, et al.
Published: (2024)
Mean-field Chaos Diffusion Models
by: Park, Sungwoo, et al.
Published: (2024)
by: Park, Sungwoo, et al.
Published: (2024)
CLIP-KOA: Enhancing Knee Osteoarthritis Diagnosis with Multi-Modal Learning and Symmetry-Aware Loss Functions
by: Jeong, Yejin, et al.
Published: (2025)
by: Jeong, Yejin, et al.
Published: (2025)
An intuitive multi-frequency feature representation for SO(3)-equivariant networks
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
FixCLR: Negative-Class Contrastive Learning for Semi-Supervised Domain Generalization
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
by: Kim, Jiyeong, et al.
Published: (2026)
by: Kim, Jiyeong, et al.
Published: (2026)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
Efficient Odd-One-Out Anomaly Detection
by: Chito, Silvio, et al.
Published: (2025)
by: Chito, Silvio, et al.
Published: (2025)
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
by: Kwon, JuneHyoung, et al.
Published: (2026)
by: Kwon, JuneHyoung, et al.
Published: (2026)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
by: Min, Jeongho, et al.
Published: (2025)
by: Min, Jeongho, et al.
Published: (2025)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
by: Li, Xiujun, et al.
Published: (2023)
by: Li, Xiujun, et al.
Published: (2023)
DiffInject: Revisiting Debias via Synthetic Data Generation using Diffusion-based Style Injection
by: Ko, Donggeun, et al.
Published: (2024)
by: Ko, Donggeun, et al.
Published: (2024)
Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
by: Son, Moo Hyun, et al.
Published: (2025)
by: Son, Moo Hyun, et al.
Published: (2025)
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
by: Kwon, Yejin, et al.
Published: (2025)
by: Kwon, Yejin, et al.
Published: (2025)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
Odd-One-Out: Anomaly Detection by Comparing with Neighbors
by: Bhunia, Ankan, et al.
Published: (2024)
by: Bhunia, Ankan, et al.
Published: (2024)
Open Set Recognition for Endoscopic Image Classification: A Deep Learning Approach on the Kvasir Dataset
by: Moazzami, Kasra, et al.
Published: (2025)
by: Moazzami, Kasra, et al.
Published: (2025)
Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image Segmentation
by: Yu, Seonghoon, et al.
Published: (2024)
by: Yu, Seonghoon, et al.
Published: (2024)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
by: La, Tuan-Vinh, et al.
Published: (2025)
by: La, Tuan-Vinh, et al.
Published: (2025)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
by: Jang, You-Won, et al.
Published: (2025)
by: Jang, You-Won, et al.
Published: (2025)
AoP-SAM: Automation of Prompts for Efficient Segmentation
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
Similar Items
-
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024) -
An Ordinary Differential Equation Sampler with Stochastic Start for Diffusion Bridge Models
by: Wang, Yuang, et al.
Published: (2024) -
NODER: Image Sequence Regression Based on Neural Ordinary Differential Equations
by: Bai, Hao, et al.
Published: (2024) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023) -
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
by: Yu, Seonghoon, et al.
Published: (2025)