Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yu, Zhang, Yuxin, Cao, Juan, Gao, Lin, Wang, Chunyu, Deussen, Oliver, Lee, Tong-Yee, Tang, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation
by: Luo, Minyan, et al.
Published: (2026)
by: Luo, Minyan, et al.
Published: (2026)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
by: Xu, Yu, et al.
Published: (2025)
by: Xu, Yu, et al.
Published: (2025)
Dance-to-Music Generation with Encoder-based Textual Inversion
by: Li, Sifei, et al.
Published: (2024)
by: Li, Sifei, et al.
Published: (2024)
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
by: Li, Sifei, et al.
Published: (2026)
by: Li, Sifei, et al.
Published: (2026)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Interactive Visual Assessment for Text-to-Image Generation Models
by: Mi, Xiaoyue, et al.
Published: (2024)
by: Mi, Xiaoyue, et al.
Published: (2024)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025)
by: Fang, Haipeng, et al.
Published: (2025)
DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer
by: Huang, Nisha, et al.
Published: (2026)
by: Huang, Nisha, et al.
Published: (2026)
Towards Agentic Schema Refinement
by: Rissaki, Agapi, et al.
Published: (2024)
by: Rissaki, Agapi, et al.
Published: (2024)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
UADAPy: An Uncertainty-Aware Visualization and Analysis Toolbox
by: Paetzold, Patrick, et al.
Published: (2024)
by: Paetzold, Patrick, et al.
Published: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
by: Zhou, Chenming, et al.
Published: (2025)
by: Zhou, Chenming, et al.
Published: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
Beyond Correctness: Learning Robust Reasoning via Transfer
by: Lee, Hyunseok, et al.
Published: (2026)
by: Lee, Hyunseok, et al.
Published: (2026)
From Web to Pixels: Bringing Agentic Search into Visual Perception
by: Yang, Bokang, et al.
Published: (2026)
by: Yang, Bokang, et al.
Published: (2026)
SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation
by: Radanliev, Petar, et al.
Published: (2026)
by: Radanliev, Petar, et al.
Published: (2026)
ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning
by: Liu, Shifeng, et al.
Published: (2026)
by: Liu, Shifeng, et al.
Published: (2026)
MetaphorVU: Towards Metaphorical Video Understanding
by: Li, Zhuoqun, et al.
Published: (2026)
by: Li, Zhuoqun, et al.
Published: (2026)
Comparing Design Metaphors and User-Driven Metaphors for Interaction Design
by: Bullock, Beleicia, et al.
Published: (2026)
by: Bullock, Beleicia, et al.
Published: (2026)
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
by: Ning, Zhiwei, et al.
Published: (2026)
by: Ning, Zhiwei, et al.
Published: (2026)
Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework
by: Huang, Ziyao, et al.
Published: (2024)
by: Huang, Ziyao, et al.
Published: (2024)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
Schema-Driven Information Extraction from Heterogeneous Tables
by: Bai, Fan, et al.
Published: (2023)
by: Bai, Fan, et al.
Published: (2023)
MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Na'vi or Knave: Jailbreaking Language Models via Metaphorical Avatars
by: Yan, Yu, et al.
Published: (2024)
by: Yan, Yu, et al.
Published: (2024)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
by: Xiang, Wenzhao, et al.
Published: (2026)
by: Xiang, Wenzhao, et al.
Published: (2026)
XML Schema Languages: Beyond DTD.
by: Ioannides, Demetrios
Published: (2000)
by: Ioannides, Demetrios
Published: (2000)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
by: Shen, Yucheng, et al.
Published: (2026)
by: Shen, Yucheng, et al.
Published: (2026)
Beyond Visual Appearances: Privacy-sensitive Objects Identification via Hybrid Graph Reasoning
by: Jiang, Zhuohang, et al.
Published: (2024)
by: Jiang, Zhuohang, et al.
Published: (2024)
Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
by: Han, Tianyang, et al.
Published: (2025)
by: Han, Tianyang, et al.
Published: (2025)
DIAL-KG: Schema-Free Incremental Knowledge Graph Construction via Dynamic Schema Induction and Evolution-Intent Assessment
by: Bao, Weidong, et al.
Published: (2026)
by: Bao, Weidong, et al.
Published: (2026)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
by: Chng, Yong Xien, et al.
Published: (2025)
by: Chng, Yong Xien, et al.
Published: (2025)
Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
Similar Items
-
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation
by: Luo, Minyan, et al.
Published: (2026) -
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
by: Xu, Yu, et al.
Published: (2024) -
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
by: Xu, Yu, et al.
Published: (2024) -
In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
by: Xu, Yu, et al.
Published: (2025) -
Dance-to-Music Generation with Encoder-based Textual Inversion
by: Li, Sifei, et al.
Published: (2024)