Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chenhao, Niu, Yazhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
by: Zhang, Chenhao, et al.
Published: (2024)
by: Zhang, Chenhao, et al.
Published: (2024)
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
by: Du, Chengyi, et al.
Published: (2026)
by: Du, Chengyi, et al.
Published: (2026)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
by: Bai, Yatong, et al.
Published: (2024)
by: Bai, Yatong, et al.
Published: (2024)
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
Computer Vision for Multimedia Geolocation in Human Trafficking Investigation: A Systematic Literature Review
by: Bamigbade, Opeyemi, et al.
Published: (2024)
by: Bamigbade, Opeyemi, et al.
Published: (2024)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
Position: Towards Implicit Prompt For Text-To-Image Models
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
An AI-Enabled Framework Within Reach for Enhancing Healthcare Sustainability and Fairness
by: Huang, Bin, et al.
Published: (2024)
by: Huang, Bin, et al.
Published: (2024)
From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
by: Cupini, Paolo, et al.
Published: (2026)
by: Cupini, Paolo, et al.
Published: (2026)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
by: Liu, Ziqiang, et al.
Published: (2024)
by: Liu, Ziqiang, et al.
Published: (2024)
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
by: Hashmi, Ammarah, et al.
Published: (2024)
by: Hashmi, Ammarah, et al.
Published: (2024)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
Unsupervised Deep Learning Image Verification Method
by: Solomon, Enoch, et al.
Published: (2023)
by: Solomon, Enoch, et al.
Published: (2023)
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
by: Gritsevskiy, Andrew, et al.
Published: (2024)
by: Gritsevskiy, Andrew, et al.
Published: (2024)
Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition
by: Akben, Mustafa, et al.
Published: (2025)
by: Akben, Mustafa, et al.
Published: (2025)
Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?
by: Malekzadeh, Milad, et al.
Published: (2025)
by: Malekzadeh, Milad, et al.
Published: (2025)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
INFELM: In-depth Fairness Evaluation of Large Text-To-Image Models
by: Jin, Di, et al.
Published: (2024)
by: Jin, Di, et al.
Published: (2024)
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025)
by: Deng, Boyang, et al.
Published: (2025)
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
by: Shi, Chuancheng, et al.
Published: (2025)
by: Shi, Chuancheng, et al.
Published: (2025)
Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification
by: Doh, Miriam, et al.
Published: (2026)
by: Doh, Miriam, et al.
Published: (2026)
Towards Understanding Unsafe Video Generation
by: Pang, Yan, et al.
Published: (2024)
by: Pang, Yan, et al.
Published: (2024)
Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face Detection
by: Hu, Juan, et al.
Published: (2025)
by: Hu, Juan, et al.
Published: (2025)
FairREAD: Re-fusing Demographic Attributes after Disentanglement for Fair Medical Image Classification
by: Gao, Yicheng, et al.
Published: (2024)
by: Gao, Yicheng, et al.
Published: (2024)
Smiling Women Pitching Down: Auditing Representational and Presentational Gender Biases in Image Generative AI
by: Sun, Luhang, et al.
Published: (2023)
by: Sun, Luhang, et al.
Published: (2023)
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
by: Chen, Xinwang, et al.
Published: (2024)
by: Chen, Xinwang, et al.
Published: (2024)
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification
by: Yu, Xinyao, et al.
Published: (2024)
by: Yu, Xinyao, et al.
Published: (2024)
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
by: Tan, Kaizhen, et al.
Published: (2025)
by: Tan, Kaizhen, et al.
Published: (2025)
Auditing Gender Presentation Differences in Text-to-Image Models
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
DreamPainter: Image Background Inpainting for E-commerce Scenarios
by: Zhao, Sijie, et al.
Published: (2025)
by: Zhao, Sijie, et al.
Published: (2025)
Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
by: Dong, Shuai, et al.
Published: (2025)
by: Dong, Shuai, et al.
Published: (2025)
Cycle-YOLO: A Efficient and Robust Framework for Pavement Damage Detection
by: Li, Zhengji, et al.
Published: (2024)
by: Li, Zhengji, et al.
Published: (2024)
Fair Text-to-Image Diffusion via Fair Mapping
by: Li, Jia, et al.
Published: (2023)
by: Li, Jia, et al.
Published: (2023)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
by: Qiu, Chenhao, et al.
Published: (2026)
by: Qiu, Chenhao, et al.
Published: (2026)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
Similar Items
-
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026) -
Can MLLMs Understand the Deep Implication Behind Chinese Images?
by: Zhang, Chenhao, et al.
Published: (2024) -
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
by: Du, Chengyi, et al.
Published: (2026) -
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026) -
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
by: Bai, Yatong, et al.
Published: (2024)