Pixels, Patterns, but No Poetry: To See The World like Humans
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Hongcheng, Huang, Zihao, Xu, Lin, Tang, Jingyi, Li, Xinhao, Liu, Yue, Li, Haoyang, Hu, Taihang, Lin, Minhua, Yang, Xinlong, Wu, Ge, Bi, Balong, Chen, Hongyu, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024)
by: Gao, Hongcheng, et al.
Published: (2024)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
by: Long, Lin, et al.
Published: (2025)
by: Long, Lin, et al.
Published: (2025)
Open 3D World in Autonomous Driving
by: Cheng, Xinlong, et al.
Published: (2024)
by: Cheng, Xinlong, et al.
Published: (2024)
PEAR: Pixel-aligned Expressive humAn mesh Recovery
by: Wu, Jiahao, et al.
Published: (2026)
by: Wu, Jiahao, et al.
Published: (2026)
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)
by: Xu, Gangwei, et al.
Published: (2026)
Diffusion Feedback Helps CLIP See Better
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
by: Wang, Zongwei, et al.
Published: (2026)
by: Wang, Zongwei, et al.
Published: (2026)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
by: Tang, Qi, et al.
Published: (2024)
by: Tang, Qi, et al.
Published: (2024)
TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
by: Lu, Jiahao, et al.
Published: (2025)
by: Lu, Jiahao, et al.
Published: (2025)
Can AI Write Classical Chinese Poetry like Humans? An Empirical Study Inspired by Turing Test
by: Deng, Zekun, et al.
Published: (2024)
by: Deng, Zekun, et al.
Published: (2024)
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
by: Jamil, Sofia, et al.
Published: (2025)
by: Jamil, Sofia, et al.
Published: (2025)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
by: Liu, Feiyang, et al.
Published: (2024)
by: Liu, Feiyang, et al.
Published: (2024)
Pixelis: Reasoning in Pixels, from Seeing to Acting
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
by: Tang, Jingyi, et al.
Published: (2026)
by: Tang, Jingyi, et al.
Published: (2026)
OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2024)
by: Xia, Zhongyu, et al.
Published: (2024)
AMPLE: Emotion-Aware Multimodal Fusion Prompt Learning for Fake News Detection
by: Xu, Xiaoman, et al.
Published: (2024)
by: Xu, Xiaoman, et al.
Published: (2024)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
Seeing World Dynamics in a Nutshell
by: Shen, Qiuhong, et al.
Published: (2025)
by: Shen, Qiuhong, et al.
Published: (2025)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
Theoretical and Empirical Validation of Heston Model
by: Cao, Zheng, et al.
Published: (2024)
by: Cao, Zheng, et al.
Published: (2024)
Event-Based Method for High-Speed 3D Deformation Measurement under Extreme Illumination Conditions
by: Guan, Banglei, et al.
Published: (2026)
by: Guan, Banglei, et al.
Published: (2026)
PixelsDB: Serverless and NL-Aided Data Analytics with Flexible Service Levels and Prices
by: Bian, Haoqiong, et al.
Published: (2024)
by: Bian, Haoqiong, et al.
Published: (2024)
IntegratedPIFu: Integrated Pixel Aligned Implicit Function for Single-view Human Reconstruction
by: Chan, Kennard Yanting, et al.
Published: (2022)
by: Chan, Kennard Yanting, et al.
Published: (2022)
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions
by: Guo, Zhongbin, et al.
Published: (2026)
by: Guo, Zhongbin, et al.
Published: (2026)
Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
by: Ren, Yudan, et al.
Published: (2025)
by: Ren, Yudan, et al.
Published: (2025)
Towards Immersive Mixed Reality Street Play: Understanding Co-located Bodily Play with See-through Head-mounted Displays in Public Spaces
by: Hu, Botao Amber, et al.
Published: (2025)
by: Hu, Botao Amber, et al.
Published: (2025)
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
KPoEM: A Human-Annotated Dataset for Emotion Classification and RAG-Based Poetry Generation in Korean Modern Poetry
by: Lim, Iro, et al.
Published: (2025)
by: Lim, Iro, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
by: Zeng, Bohan, et al.
Published: (2026)
by: Zeng, Bohan, et al.
Published: (2026)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
by: Ma, Baorui, et al.
Published: (2024)
by: Ma, Baorui, et al.
Published: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
by: Yuan, Qianhao, et al.
Published: (2026)
by: Yuan, Qianhao, et al.
Published: (2026)
Similar Items
-
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025) -
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024) -
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025) -
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
by: Long, Lin, et al.
Published: (2025) -
Open 3D World in Autonomous Driving
by: Cheng, Xinlong, et al.
Published: (2024)