Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Bumsoo, Shin, Wonseop, Lee, Kyuchul, Jung, Yonghoon, Seo, Sanghyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024)
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
von: Kang, YoonJe, et al.
Veröffentlicht: (2025)
von: Kang, YoonJe, et al.
Veröffentlicht: (2025)
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
von: Muqeet, Abdul, et al.
Veröffentlicht: (2023)
von: Muqeet, Abdul, et al.
Veröffentlicht: (2023)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
von: Fan, Xinqi, et al.
Veröffentlicht: (2025)
von: Fan, Xinqi, et al.
Veröffentlicht: (2025)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
LPM 1.0: Video-based Character Performance Model
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
SAFIRE: Segment Any Forged Image Region
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
Audio Visual Segmentation Through Text Embeddings
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
Detection and Recovery of Adversarial Slow-Pose Drift in Offloaded Visual-Inertial Odometry
von: Saha, Soruya, et al.
Veröffentlicht: (2025)
von: Saha, Soruya, et al.
Veröffentlicht: (2025)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
A Simple Baseline with Single-encoder for Referring Image Segmentation
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
von: Kim, Geewook, et al.
Veröffentlicht: (2025)
von: Kim, Geewook, et al.
Veröffentlicht: (2025)
DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024) -
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
von: Kim, Bumsoo, et al.
Veröffentlicht: (2024) -
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
von: Kang, YoonJe, et al.
Veröffentlicht: (2025) -
Video Face Re-Aging: Toward Temporally Consistent Face Re-Aging
von: Muqeet, Abdul, et al.
Veröffentlicht: (2023) -
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
von: Zheng, Jiayi, et al.
Veröffentlicht: (2025)