TV2TV: A Unified Framework for Interleaved Language and Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Xiaochuang, Emad, Youssef, Hall, Melissa, Nguyen, John, Padthe, Karthik, Robbins, Liam, Bar, Amir, Chen, Delong, Drozdzal, Michal, Elbayad, Maha, Hu, Yushi, Li, Shang-Wen, Roy, Sreya Dutta, Verbeek, Jakob, Wang, XuDong, Ghazvininejad, Marjan, Zettlemoyer, Luke, Dinan, Emily |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
by: Hu, Yushi, et al.
Published: (2025)
by: Hu, Yushi, et al.
Published: (2025)
Unified Text-Image Generation with Weakness-Targeted Post-Training
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
Text-Guided Semantic Image Encoder
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
by: Han, Xiaochuang, et al.
Published: (2023)
by: Han, Xiaochuang, et al.
Published: (2023)
VUGEN: Visual Understanding priors for GENeration
by: Chen, Xiangyi, et al.
Published: (2025)
by: Chen, Xiangyi, et al.
Published: (2025)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
by: Han, Xiaochuang, et al.
Published: (2024)
by: Han, Xiaochuang, et al.
Published: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
by: Kamath, Amita, et al.
Published: (2025)
by: Kamath, Amita, et al.
Published: (2025)
Merging Text Transformer Models from Different Initializations
by: Verma, Neha, et al.
Published: (2024)
by: Verma, Neha, et al.
Published: (2024)
Self-Improving VLM Judges Without Human Annotations
by: Lin, Inna Wanyin, et al.
Published: (2025)
by: Lin, Inna Wanyin, et al.
Published: (2025)
ALMA: Alignment with Minimal Annotation
by: Yasunaga, Michihiro, et al.
Published: (2024)
by: Yasunaga, Michihiro, et al.
Published: (2024)
Entropy Rectifying Guidance for Diffusion and Flow Models
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
TV PUBLICA VERSUS TV PRIVADA
by: ROMANO, VICENTE
Published: (2007)
by: ROMANO, VICENTE
Published: (2007)
Increasing the Utility of Synthetic Images through Chamfer Guidance
by: Dall'Asen, Nicola, et al.
Published: (2025)
by: Dall'Asen, Nicola, et al.
Published: (2025)
Consistency-diversity-realism Pareto fronts of conditional image generative models
by: Astolfi, Pietro, et al.
Published: (2024)
by: Astolfi, Pietro, et al.
Published: (2024)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
What Should We Watch Today? The Role of Shared TV Watching in Couples' Daily Shared Reality
by: Ravid Haruvi, et al.
Published: (2026)
by: Ravid Haruvi, et al.
Published: (2026)
TV Series
Published: (2017)
Published: (2017)
Treffpunkt TV
Published: (2024)
Published: (2024)
Treffpunkt TV
Published: (2024)
Published: (2024)
RACE AND REPRESENTATION ON TV: THE INFLUENCE OF TV STATUS ON LATINO IDENTITIES
by: Jesus Augusto Gonzalez
Published: (2017)
by: Jesus Augusto Gonzalez
Published: (2017)
Representation Deficiency in Masked Language Modeling
by: Meng, Yu, et al.
Published: (2023)
by: Meng, Yu, et al.
Published: (2023)
Improving Factuality with Explicit Working Memory
by: Chen, Mingda, et al.
Published: (2024)
by: Chen, Mingda, et al.
Published: (2024)
Boosting Latent Diffusion with Perceptual Objectives
by: Berrada, Tariq, et al.
Published: (2024)
by: Berrada, Tariq, et al.
Published: (2024)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
by: Wang, Dong, et al.
Published: (2025)
by: Wang, Dong, et al.
Published: (2025)
Locally Correct Interleavings between Merge Trees
by: Beurskens, Thijs, et al.
Published: (2025)
by: Beurskens, Thijs, et al.
Published: (2025)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Arab TV-Audiences
Published: (2020)
Published: (2020)
Arab TV-Audiences
Published: (2020)
Published: (2020)
Television before TV
by: Weber, Anne-Katrin
Published: (2022)
by: Weber, Anne-Katrin
Published: (2022)
Analyzing Reality TV
by: Lois Elfman
Published: (2025)
by: Lois Elfman
Published: (2025)
Similar Items
-
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
by: Hu, Yushi, et al.
Published: (2025) -
Unified Text-Image Generation with Weakness-Targeted Post-Training
by: Chen, Jiahui, et al.
Published: (2026) -
Text-Guided Semantic Image Encoder
by: Thirukovalluru, Raghuveer, et al.
Published: (2025) -
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025) -
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
by: Han, Xiaochuang, et al.
Published: (2023)