Visually Dehallucinative Instruction Generation: Know What You Don't Know
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cha, Sungguk, Lee, Jusung, Lee, Younghyun, Yang, Cheoljong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visually Dehallucinative Instruction Generation
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2024)
von: Li, Senmao, et al.
Veröffentlicht: (2024)
Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
von: Thozhiyoor, Varun Varma, et al.
Veröffentlicht: (2025)
von: Thozhiyoor, Varun Varma, et al.
Veröffentlicht: (2025)
Specifying What You Know or Not for Multi-Label Class-Incremental Learning
von: Zhang, Aoting, et al.
Veröffentlicht: (2025)
von: Zhang, Aoting, et al.
Veröffentlicht: (2025)
I Detect What I Don't Know: Incremental Anomaly Learning with Stochastic Weight Averaging-Gaussian for Oracle-Free Medical Imaging
von: Yadav, Nand Kumar, et al.
Veröffentlicht: (2025)
von: Yadav, Nand Kumar, et al.
Veröffentlicht: (2025)
Know What You do Not Know: Verbalized Uncertainty Estimation Robustness on Corrupted Images in Vision-Language Models
von: Borszukovszki, Mirko, et al.
Veröffentlicht: (2025)
von: Borszukovszki, Mirko, et al.
Veröffentlicht: (2025)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
von: Wu, Tao, et al.
Veröffentlicht: (2025)
von: Wu, Tao, et al.
Veröffentlicht: (2025)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2025)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2025)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
von: Kwon, Mincheol, et al.
Veröffentlicht: (2026)
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
All You Need to Know About Training Image Retrieval Models
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
von: Berton, Gabriele, et al.
Veröffentlicht: (2025)
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
von: Jin, Er, et al.
Veröffentlicht: (2025)
von: Jin, Er, et al.
Veröffentlicht: (2025)
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
von: Du, Xiaodan, et al.
Veröffentlicht: (2023)
von: Du, Xiaodan, et al.
Veröffentlicht: (2023)
Don't Judge Before You CLIP: A Unified Approach for Perceptual Tasks
von: Zalcher, Amit, et al.
Veröffentlicht: (2025)
von: Zalcher, Amit, et al.
Veröffentlicht: (2025)
Point What You Mean: Visually Grounded Instruction Policy
von: Yu, Hang, et al.
Veröffentlicht: (2025)
von: Yu, Hang, et al.
Veröffentlicht: (2025)
Text Embedding Knows How to Quantize Text-Guided Diffusion Models
von: Lee, Hongjae, et al.
Veröffentlicht: (2025)
von: Lee, Hongjae, et al.
Veröffentlicht: (2025)
VisKnow: Constructing Visual Knowledge Base for Object Understanding
von: Yao, Ziwei, et al.
Veröffentlicht: (2025)
von: Yao, Ziwei, et al.
Veröffentlicht: (2025)
Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
von: Lin, Kaiqing, et al.
Veröffentlicht: (2025)
von: Lin, Kaiqing, et al.
Veröffentlicht: (2025)
PTQ4VM: Post-Training Quantization for Visual Mamba
von: Cho, Younghyun, et al.
Veröffentlicht: (2024)
von: Cho, Younghyun, et al.
Veröffentlicht: (2024)
"I Know It When I See It": Mood Spaces for Connecting and Expressing Visual Concepts
von: Yang, Huzheng, et al.
Veröffentlicht: (2025)
von: Yang, Huzheng, et al.
Veröffentlicht: (2025)
ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
von: Egin, Anil, et al.
Veröffentlicht: (2026)
von: Egin, Anil, et al.
Veröffentlicht: (2026)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
If At First You Don't Succeed: Test Time Re-ranking for Zero-shot, Cross-domain Retrieval
von: Hudson, Finlay G. C., et al.
Veröffentlicht: (2023)
von: Hudson, Finlay G. C., et al.
Veröffentlicht: (2023)
Discovering and Mitigating Visual Biases through Keyword Explanation
von: Kim, Younghyun, et al.
Veröffentlicht: (2023)
von: Kim, Younghyun, et al.
Veröffentlicht: (2023)
Explorations of the Softmax Space: Knowing When the Neural Network Doesn't Know
von: Sikar, Daniel, et al.
Veröffentlicht: (2025)
von: Sikar, Daniel, et al.
Veröffentlicht: (2025)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
von: Zou, Xin, et al.
Veröffentlicht: (2025)
von: Zou, Xin, et al.
Veröffentlicht: (2025)
ProCreate, Don't Reproduce! Propulsive Energy Diffusion for Creative Generation
von: Lu, Jack, et al.
Veröffentlicht: (2024)
von: Lu, Jack, et al.
Veröffentlicht: (2024)
Don't Play Favorites: Minority Guidance for Diffusion Models
von: Um, Soobin, et al.
Veröffentlicht: (2023)
von: Um, Soobin, et al.
Veröffentlicht: (2023)
Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
von: Jiang, Tianchong, et al.
Veröffentlicht: (2025)
von: Jiang, Tianchong, et al.
Veröffentlicht: (2025)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
von: Hwang, Geunmin, et al.
Veröffentlicht: (2025)
von: Hwang, Geunmin, et al.
Veröffentlicht: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
von: Baid, Ami, et al.
Veröffentlicht: (2026)
von: Baid, Ami, et al.
Veröffentlicht: (2026)
Beyond Words: Multimodal LLM Knows When to Speak
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Visually Dehallucinative Instruction Generation
von: Cha, Sungguk, et al.
Veröffentlicht: (2024) -
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024) -
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
von: Mei, Zhiting, et al.
Veröffentlicht: (2025) -
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024) -
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2024)