Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miyai, Atsuyuki, Yang, Jingkang, Zhang, Jingyang, Ming, Yifei, Lin, Yueqian, Yu, Qing, Irie, Go, Joty, Shafiq, Li, Yixuan, Li, Hai, Liu, Ziwei, Yamasaki, Toshihiko, Aizawa, Kiyoharu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
von: Noda, Shiho, et al.
Veröffentlicht: (2025)
von: Noda, Shiho, et al.
Veröffentlicht: (2025)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2026)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Generalized Out-of-Distribution Detection: A Survey
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
von: Yang, Jingkang, et al.
Veröffentlicht: (2021)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
OpenOOD v1.5: Enhanced Benchmark for Out-of-Distribution Detection
von: Zhang, Jingyang, et al.
Veröffentlicht: (2023)
von: Zhang, Jingyang, et al.
Veröffentlicht: (2023)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
NAACL2025 Tutorial: Adaptation of Large Language Models
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
von: Ming, Yifei, et al.
Veröffentlicht: (2023)
von: Ming, Yifei, et al.
Veröffentlicht: (2023)
Harnessing LLM Agents with Skill Programs
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
von: Liu, Hongjun, et al.
Veröffentlicht: (2026)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Manga109Dialog: A Large-scale Dialogue Dataset for Comics Speaker Detection
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
von: Li, Yingxuan, et al.
Veröffentlicht: (2023)
SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
von: Lin, Yueqian, et al.
Veröffentlicht: (2023)
von: Lin, Yueqian, et al.
Veröffentlicht: (2023)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
von: Zhao, Zaiying, et al.
Veröffentlicht: (2025)
von: Zhao, Zaiying, et al.
Veröffentlicht: (2025)
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
von: Tanji, Naoto, et al.
Veröffentlicht: (2025)
von: Tanji, Naoto, et al.
Veröffentlicht: (2025)
Probing Prompt Design for Socially Compliant Robot Navigation with Vision Language Models
von: Xiao, Ling, et al.
Veröffentlicht: (2026)
von: Xiao, Ling, et al.
Veröffentlicht: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
Investigating the Perception of Facial Anonymization Techniques in 360° Videos
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
von: Otonari, Takashi, et al.
Veröffentlicht: (2024)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
Privacy Protection and Video Manipulation in Immersive Media
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
von: Wöhler, Leslie, et al.
Veröffentlicht: (2024)
Guided Image Synthesis via Initial Image Editing in Diffusion Model
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
von: Mao, Jiafeng, et al.
Veröffentlicht: (2023)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
von: Toyooka, Mashiro, et al.
Veröffentlicht: (2025)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
von: Asano, Shunta, et al.
Veröffentlicht: (2026)
von: Asano, Shunta, et al.
Veröffentlicht: (2026)
HYPO: Hyperspherical Out-of-Distribution Generalization
von: Bai, Haoyue, et al.
Veröffentlicht: (2024)
von: Bai, Haoyue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2023) -
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024) -
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
von: Noda, Shiho, et al.
Veröffentlicht: (2025) -
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025) -
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)