Transfer between Modalities with MetaQueries
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Xichen, Shukla, Satya Narayan, Singh, Aashu, Zhao, Zhuokai, Mishra, Shlok Kumar, Wang, Jialiang, Xu, Zhiyang, Chen, Jiuhai, Li, Kunpeng, Juefei-Xu, Felix, Hou, Ji, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
Think Then Embed: Generative Context Improves Multimodal Embedding
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
Heteroscedastic Temporal Variational Autoencoder For Irregular Time Series
by: Shukla, Satya Narayan, et al.
Published: (2021)
by: Shukla, Satya Narayan, et al.
Published: (2021)
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Depicting Culture and Identity in Amitav Ghosh’s The Shadow Lines
by: Satya Narayan
Published: (2017)
by: Satya Narayan
Published: (2017)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
An Attribute-Based Measure of Video Complexity
by: Sarkar, Aditya, et al.
Published: (2026)
by: Sarkar, Aditya, et al.
Published: (2026)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
by: Li, Chenyu, et al.
Published: (2025)
by: Li, Chenyu, et al.
Published: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
by: Kodaira, Akio, et al.
Published: (2025)
by: Kodaira, Akio, et al.
Published: (2025)
Cambrian-P: Pose-Grounded Video Understanding
by: Yang, Jihan, et al.
Published: (2026)
by: Yang, Jihan, et al.
Published: (2026)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
BLIP3o-NEXT: Next Frontier of Native Image Generation
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
by: Xu, Zhiyang, et al.
Published: (2025)
by: Xu, Zhiyang, et al.
Published: (2025)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
MoCha: Towards Movie-Grade Talking Character Synthesis
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
Xray-Visual Models: Scaling Vision models on Industry Scale Data
by: Mishra, Shlok, et al.
Published: (2026)
by: Mishra, Shlok, et al.
Published: (2026)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
by: Wang, Hongjie, et al.
Published: (2024)
by: Wang, Hongjie, et al.
Published: (2024)
Compositional Visual Planning via Inference-Time Diffusion Scaling
by: Zhang, Yixin, et al.
Published: (2026)
by: Zhang, Yixin, et al.
Published: (2026)
AV$_3$Sb$_5$ kagome superconductors: a review with transport measurements
by: Xu, Zhuokai, et al.
Published: (2025)
by: Xu, Zhuokai, et al.
Published: (2025)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake
by: Yang, Di, et al.
Published: (2024)
by: Yang, Di, et al.
Published: (2024)
Query Lower Bounds for Diffusion Sampling
by: Xun, Zhiyang, et al.
Published: (2026)
by: Xun, Zhiyang, et al.
Published: (2026)
Automated Data Curation for Robust Language Model Fine-Tuning
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
On the Query Complexity of Training Data Reconstruction in Private Learning
by: Mukherjee, Prateeti, et al.
Published: (2023)
by: Mukherjee, Prateeti, et al.
Published: (2023)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
by: Zhao, Shu, et al.
Published: (2025)
by: Zhao, Shu, et al.
Published: (2025)
LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
by: Shen, Ying, et al.
Published: (2025)
by: Shen, Ying, et al.
Published: (2025)
Multimodal Guidance Network for Missing-Modality Inference in Content Moderation
by: Zhao, Zhuokai, et al.
Published: (2023)
by: Zhao, Zhuokai, et al.
Published: (2023)
Similar Items
-
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025) -
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025) -
Think Then Embed: Generative Context Improves Multimodal Embedding
by: Cui, Xuanming, et al.
Published: (2025) -
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025) -
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)