Salvato in:
| Autori principali: | Chen, Haoyu, Liu, Qing, Zhou, Yuqian, Zhang, He, Wang, Zhaowen, Ren, Mengwei, Ren, Jingjing, Wang, Xiang, Lin, Zhe, Zhu, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.07540 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
di: Yang, Jinrui, et al.
Pubblicazione: (2026)
di: Yang, Jinrui, et al.
Pubblicazione: (2026)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
di: Ju, Xuan, et al.
Pubblicazione: (2025)
di: Ju, Xuan, et al.
Pubblicazione: (2025)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
di: Liu, Jinkun, et al.
Pubblicazione: (2026)
di: Liu, Jinkun, et al.
Pubblicazione: (2026)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
di: Li, Aaron Branson Cigres, et al.
Pubblicazione: (2026)
di: Li, Aaron Branson Cigres, et al.
Pubblicazione: (2026)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
di: Zhang, Yabo, et al.
Pubblicazione: (2026)
di: Zhang, Yabo, et al.
Pubblicazione: (2026)
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
di: Zhang, Yuxiang, et al.
Pubblicazione: (2025)
di: Zhang, Yuxiang, et al.
Pubblicazione: (2025)
Generative Image Layer Decomposition with Visual Effects
di: Yang, Jinrui, et al.
Pubblicazione: (2024)
di: Yang, Jinrui, et al.
Pubblicazione: (2024)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
di: Wang, Bingli, et al.
Pubblicazione: (2026)
di: Wang, Bingli, et al.
Pubblicazione: (2026)
HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
di: Wang, Xiang, et al.
Pubblicazione: (2025)
di: Wang, Xiang, et al.
Pubblicazione: (2025)
LoViC: Efficient Long Video Generation with Context Compression
di: Jiang, Jiaxiu, et al.
Pubblicazione: (2025)
di: Jiang, Jiaxiu, et al.
Pubblicazione: (2025)
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
di: Cai, Yuanhao, et al.
Pubblicazione: (2024)
di: Cai, Yuanhao, et al.
Pubblicazione: (2024)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
di: Li, Qingyun, et al.
Pubblicazione: (2024)
di: Li, Qingyun, et al.
Pubblicazione: (2024)
Sulfurized Polyacrylonitrile Cathodes With Rapid Redox Kinetics for High‐Capacity and Long‐Cycle‐Life Lithium‐Sulfur Batteries
di: Liang Tian, et al.
Pubblicazione: (2025)
di: Liang Tian, et al.
Pubblicazione: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)
di: Tian, Changyao, et al.
Pubblicazione: (2024)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
di: Hu, Yucheng, et al.
Pubblicazione: (2026)
di: Hu, Yucheng, et al.
Pubblicazione: (2026)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
di: Wan, Guangya, et al.
Pubblicazione: (2025)
di: Wan, Guangya, et al.
Pubblicazione: (2025)
McCast: Memory-Guided Latent Drift Correction for Long-Horizon Precipitation Nowcasting
di: Wen, Penghui, et al.
Pubblicazione: (2026)
di: Wen, Penghui, et al.
Pubblicazione: (2026)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
di: Chen, Junyi, et al.
Pubblicazione: (2026)
di: Chen, Junyi, et al.
Pubblicazione: (2026)
Interact-Custom: Customized Human Object Interaction Image Generation
di: Xu, Zhu, et al.
Pubblicazione: (2025)
di: Xu, Zhu, et al.
Pubblicazione: (2025)
Image Generation Diversity Issues and How to Tame Them
di: Dombrowski, Mischa, et al.
Pubblicazione: (2024)
di: Dombrowski, Mischa, et al.
Pubblicazione: (2024)
ROSA-Tuning: Enhancing Long-Context Modeling via Suffix Matching
di: Zheng, Yunao, et al.
Pubblicazione: (2026)
di: Zheng, Yunao, et al.
Pubblicazione: (2026)
Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
di: Ren, Xubin, et al.
Pubblicazione: (2025)
di: Ren, Xubin, et al.
Pubblicazione: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
di: Feng, Yukang, et al.
Pubblicazione: (2025)
di: Feng, Yukang, et al.
Pubblicazione: (2025)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
di: Nie, Ming, et al.
Pubblicazione: (2026)
di: Nie, Ming, et al.
Pubblicazione: (2026)
V-CAGE: Context-Aware Generation and Verification for Scalable Long-Horizon Embodied Tasks
di: Liu, Yaru, et al.
Pubblicazione: (2026)
di: Liu, Yaru, et al.
Pubblicazione: (2026)
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
di: Yi, Junchao, et al.
Pubblicazione: (2026)
di: Yi, Junchao, et al.
Pubblicazione: (2026)
ArbGraph: Conflict-Aware Evidence Arbitration for Reliable Long-Form Retrieval-Augmented Generation
di: Niu, Qingying, et al.
Pubblicazione: (2026)
di: Niu, Qingying, et al.
Pubblicazione: (2026)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
di: Chen, Zhekai, et al.
Pubblicazione: (2026)
di: Chen, Zhekai, et al.
Pubblicazione: (2026)
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
di: Qian, Jingjing, et al.
Pubblicazione: (2026)
di: Qian, Jingjing, et al.
Pubblicazione: (2026)
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
Interleaving Reasoning for Better Text-to-Image Generation
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents
di: Lu, Yijun, et al.
Pubblicazione: (2026)
di: Lu, Yijun, et al.
Pubblicazione: (2026)
LLM$\times$MapReduce-V2: Entropy-Driven Convolutional Test-Time Scaling for Generating Long-Form Articles from Extremely Long Resources
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
di: Jiang, Qing, et al.
Pubblicazione: (2024)
di: Jiang, Qing, et al.
Pubblicazione: (2024)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
di: Zhang, Lei, et al.
Pubblicazione: (2026)
di: Zhang, Lei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Unified and Controllable Framework for Layered Image Generation with Visual Effects
di: Yang, Jinrui, et al.
Pubblicazione: (2026) -
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
di: Ju, Xuan, et al.
Pubblicazione: (2025) -
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
di: Ren, Yiming, et al.
Pubblicazione: (2026) -
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
di: Liu, Jinkun, et al.
Pubblicazione: (2026) -
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
di: Li, Aaron Branson Cigres, et al.
Pubblicazione: (2026)