TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Bansal, Hritik, Bitton, Yonatan, Yarom, Michal, Szpektor, Idan, Grover, Aditya, Chang, Kai-Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
di: Ramos, Vasco, et al.
Pubblicazione: (2024)
di: Ramos, Vasco, et al.
Pubblicazione: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
di: Bordalo, João, et al.
Pubblicazione: (2024)
di: Bordalo, João, et al.
Pubblicazione: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
di: Zohar, Orr, et al.
Pubblicazione: (2024)
di: Zohar, Orr, et al.
Pubblicazione: (2024)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
di: Slobodkin, Aviv, et al.
Pubblicazione: (2025)
di: Slobodkin, Aviv, et al.
Pubblicazione: (2025)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
di: Gordon, Brian, et al.
Pubblicazione: (2025)
di: Gordon, Brian, et al.
Pubblicazione: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
di: Gordon, Brian, et al.
Pubblicazione: (2023)
di: Gordon, Brian, et al.
Pubblicazione: (2023)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
di: Ramos, Vasco, et al.
Pubblicazione: (2025)
di: Ramos, Vasco, et al.
Pubblicazione: (2025)
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
di: Li, Shufan, et al.
Pubblicazione: (2024)
di: Li, Shufan, et al.
Pubblicazione: (2024)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
HoneyBee: Data Recipes for Vision-Language Reasoners
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
di: Yao, Linli, et al.
Pubblicazione: (2026)
di: Yao, Linli, et al.
Pubblicazione: (2026)
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
di: Li, Shufan, et al.
Pubblicazione: (2025)
di: Li, Shufan, et al.
Pubblicazione: (2025)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
di: Zheng, Guangcong, et al.
Pubblicazione: (2025)
di: Zheng, Guangcong, et al.
Pubblicazione: (2025)
Latent Beam Diffusion Models for Generating Visual Sequences
di: Fernandes, Guilherme, et al.
Pubblicazione: (2025)
di: Fernandes, Guilherme, et al.
Pubblicazione: (2025)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
di: Wan, Yixin, et al.
Pubblicazione: (2024)
di: Wan, Yixin, et al.
Pubblicazione: (2024)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
di: Deng, Yihe, et al.
Pubblicazione: (2025)
di: Deng, Yihe, et al.
Pubblicazione: (2025)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
di: Wei, Hongchen, et al.
Pubblicazione: (2025)
di: Wei, Hongchen, et al.
Pubblicazione: (2025)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
di: Yosef, Ron, et al.
Pubblicazione: (2025)
di: Yosef, Ron, et al.
Pubblicazione: (2025)
Instance-Aligned Captions for Explainable Video Anomaly Detection
di: Song, Inpyo, et al.
Pubblicazione: (2026)
di: Song, Inpyo, et al.
Pubblicazione: (2026)
O-TALC: Steps Towards Combating Oversegmentation within Online Action Segmentation
di: Myers, Matthew Kent, et al.
Pubblicazione: (2024)
di: Myers, Matthew Kent, et al.
Pubblicazione: (2024)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
di: Rosa, Kevin Dela
Pubblicazione: (2024)
di: Rosa, Kevin Dela
Pubblicazione: (2024)
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
di: Tang, Jiyang, et al.
Pubblicazione: (2025)
di: Tang, Jiyang, et al.
Pubblicazione: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data
di: Li, Shufan, et al.
Pubblicazione: (2024)
di: Li, Shufan, et al.
Pubblicazione: (2024)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
di: Ventura, Mor, et al.
Pubblicazione: (2026)
di: Ventura, Mor, et al.
Pubblicazione: (2026)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
di: Chu, Sanghyeok, et al.
Pubblicazione: (2025)
di: Chu, Sanghyeok, et al.
Pubblicazione: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
di: Wan, Yixin, et al.
Pubblicazione: (2025)
di: Wan, Yixin, et al.
Pubblicazione: (2025)
Panoptic Captioning: An Equivalence Bridge for Image and Text
di: Lin, Kun-Yu, et al.
Pubblicazione: (2025)
di: Lin, Kun-Yu, et al.
Pubblicazione: (2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
di: Ramos, Vasco, et al.
Pubblicazione: (2024) -
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2025) -
VideoPhy: Evaluating Physical Commonsense for Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2024) -
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
di: Bordalo, João, et al.
Pubblicazione: (2024) -
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
di: Yanuka, Moran, et al.
Pubblicazione: (2024)