A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Elmoghany, Mohamed, Rossi, Ryan, Yoon, Seunghyun, Mukherjee, Subhojyoti, Bakr, Eslam, Mathur, Puneet, Wu, Gang, Lai, Viet Dac, Lipka, Nedim, Zhang, Ruiyi, Manjunatha, Varun, Nguyen, Chien, Dangi, Daksh, Salinas, Abel, Taesiri, Mohammad, Chen, Hongjie, Huang, Xiaolei, Barrow, Joe, Ahmed, Nesreen, Eldardiry, Hoda, Park, Namyong, Wang, Yu, Cho, Jaemin, Nguyen, Anh Totti, Tu, Zhengzhong, Nguyen, Thien, Manocha, Dinesh, Elhoseiny, Mohamed, Dernoncourt, Franck |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
von: Elmoghany, Mohamed, et al.
Veröffentlicht: (2026)
von: Elmoghany, Mohamed, et al.
Veröffentlicht: (2026)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
von: Giang, et al.
Veröffentlicht: (2023)
von: Giang, et al.
Veröffentlicht: (2023)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
B-score: Detecting biases in large language models using response history
von: Vo, An, et al.
Veröffentlicht: (2025)
von: Vo, An, et al.
Veröffentlicht: (2025)
KaPQA: Knowledge-Augmented Product Question-Answering
von: Eppalapally, Swetha, et al.
Veröffentlicht: (2024)
von: Eppalapally, Swetha, et al.
Veröffentlicht: (2024)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
von: Nguyen, Tin, et al.
Veröffentlicht: (2025)
von: Nguyen, Tin, et al.
Veröffentlicht: (2025)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
anguyen8/vision-llms-are-blind: official
von: Pooyan R, et al.
Veröffentlicht: (2026)
von: Pooyan R, et al.
Veröffentlicht: (2026)
Vision language models are blind: Failing to translate detailed visual features into words
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
von: Rahmanzadehgervi, Pooyan, et al.
Veröffentlicht: (2024)
Vision Language Models are Biased
von: Vo, An, et al.
Veröffentlicht: (2025)
von: Vo, An, et al.
Veröffentlicht: (2025)
ChartLens: Fine-grained Visual Attribution in Charts
von: Suri, Manan, et al.
Veröffentlicht: (2025)
von: Suri, Manan, et al.
Veröffentlicht: (2025)
Structured Uncertainty guided Clarification for LLM Agents
von: Suri, Manan, et al.
Veröffentlicht: (2025)
von: Suri, Manan, et al.
Veröffentlicht: (2025)
DynaSaur: Large Language Agents Beyond Predefined Actions
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models
von: Rawte, Vipula, et al.
Veröffentlicht: (2026)
von: Rawte, Vipula, et al.
Veröffentlicht: (2026)
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
von: Collins, Brandon, et al.
Veröffentlicht: (2026)
von: Collins, Brandon, et al.
Veröffentlicht: (2026)
Document Attribution: Examining Citation Relationships using Large Language Models
von: Rawte, Vipula, et al.
Veröffentlicht: (2025)
von: Rawte, Vipula, et al.
Veröffentlicht: (2025)
Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents
von: Qing, Peijun, et al.
Veröffentlicht: (2026)
von: Qing, Peijun, et al.
Veröffentlicht: (2026)
Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
von: Suri, Manan, et al.
Veröffentlicht: (2025)
von: Suri, Manan, et al.
Veröffentlicht: (2025)
Leveraging Habitat Information for Fine-grained Bird Identification
von: Nguyen, Tin, et al.
Veröffentlicht: (2023)
von: Nguyen, Tin, et al.
Veröffentlicht: (2023)
OATS: Opinion Aspect Target Sentiment Quadruple Extraction Dataset for Aspect-Based Sentiment Analysis
von: Chebolu, Siva Uday Sampreeth, et al.
Veröffentlicht: (2023)
von: Chebolu, Siva Uday Sampreeth, et al.
Veröffentlicht: (2023)
ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection for ABSA
von: Chebolu, Siva Uday Sampreeth, et al.
Veröffentlicht: (2024)
von: Chebolu, Siva Uday Sampreeth, et al.
Veröffentlicht: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
Congruent modular forms and anticyclotomic Iwasawa theory
von: Nguyen, Dac-Nhan-Tam
Veröffentlicht: (2025)
von: Nguyen, Dac-Nhan-Tam
Veröffentlicht: (2025)
On $λ$-invariants of congruent modular forms in the anticyclotomic, indefinite setting
von: Nguyen, Dac-Nhan-Tam
Veröffentlicht: (2025)
von: Nguyen, Dac-Nhan-Tam
Veröffentlicht: (2025)
Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
von: Fang, Jiangnan, et al.
Veröffentlicht: (2026)
von: Fang, Jiangnan, et al.
Veröffentlicht: (2026)
Sparse Personalized Text Generation with Multi-Trajectory Reasoning
von: Ni, Bo, et al.
Veröffentlicht: (2026)
von: Ni, Bo, et al.
Veröffentlicht: (2026)
Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
von: Nguyen, Giang, et al.
Veröffentlicht: (2024)
von: Nguyen, Giang, et al.
Veröffentlicht: (2024)
LongLaMP: A Benchmark for Personalized Long-form Text Generation
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
A Minkowski problem for $α$-concave functions via optimal transport
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
von: Nguyen, Hung Huy, et al.
Veröffentlicht: (2025)
von: Nguyen, Hung Huy, et al.
Veröffentlicht: (2025)
A Survey of Small Language Models
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2024)
Chitosan‐based aerogel microspheres for wastewater treatment
von: Hoang Ngoc Cuong, et al.
Veröffentlicht: (2024)
von: Hoang Ngoc Cuong, et al.
Veröffentlicht: (2024)
Lizard: An Efficient Linearization Framework for Large Language Models
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2025)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2025)
Galois cohomology of elliptic curves over anticyclotomic extensions
von: Nguyen, Dac-Nhan-Tam, et al.
Veröffentlicht: (2025)
von: Nguyen, Dac-Nhan-Tam, et al.
Veröffentlicht: (2025)
A Source Identification Problem for Subdiffusion Equations Involving Delays
von: Nguyen Van Dac, et al.
Veröffentlicht: (2025)
von: Nguyen Van Dac, et al.
Veröffentlicht: (2025)
Congruent elliptic curves over some $p$-adic Lie extensions
von: Nguyen, Dac-Nhan-Tam, et al.
Veröffentlicht: (2025)
von: Nguyen, Dac-Nhan-Tam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
von: Elmoghany, Mohamed, et al.
Veröffentlicht: (2026) -
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025) -
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
von: Giang, et al.
Veröffentlicht: (2023) -
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024) -
B-score: Detecting biases in large language models using response history
von: Vo, An, et al.
Veröffentlicht: (2025)