ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Akide, Xing, Jinbo, Mao, Chaojie, Li, Ye, Zhang, Zeyu, He, Yefei, Wang, Weijie, Wang, Zihan, Liu, Yu, Haffari, Gholamreza, Zhuang, Bohan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024)
by: Liu, Akide, et al.
Published: (2024)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
ReCA: A Parametric ReLU Composite Activation Function
by: Chidiac, John, et al.
Published: (2025)
by: Chidiac, John, et al.
Published: (2025)
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
Motion Mamba: Efficient and Long Sequence Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
by: Zhou, Junkang, et al.
Published: (2026)
by: Zhou, Junkang, et al.
Published: (2026)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
by: Huang, Shuo, et al.
Published: (2025)
by: Huang, Shuo, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
RIDE: Enhancing Large Language Model Alignment through Restyled In-Context Learning Demonstration Exemplars
by: Hua, Yuncheng, et al.
Published: (2025)
by: Hua, Yuncheng, et al.
Published: (2025)
Evidence-based Distributional Alignment for Large Language Models
by: Pham, Viet-Thanh, et al.
Published: (2026)
by: Pham, Viet-Thanh, et al.
Published: (2026)
EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
by: He, Yefei, et al.
Published: (2023)
by: He, Yefei, et al.
Published: (2023)
ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
by: You, Haochen, et al.
Published: (2025)
by: You, Haochen, et al.
Published: (2025)
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
by: Hua, Yuncheng, et al.
Published: (2024)
by: Hua, Yuncheng, et al.
Published: (2024)
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach
by: Li, Zhuang, et al.
Published: (2024)
by: Li, Zhuang, et al.
Published: (2024)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
AIPO: Learning to Reason from Active Interaction
by: Liu, Junnan, et al.
Published: (2026)
by: Liu, Junnan, et al.
Published: (2026)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models
by: Liu, Jing, et al.
Published: (2024)
by: Liu, Jing, et al.
Published: (2024)
Towards Inference-time Scaling for Continuous Space Reasoning
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
Temporal Prototyping and Hierarchical Alignment for Unsupervised Video-based Visible-Infrared Person Re-Identification
by: Li, Zhiyong, et al.
Published: (2026)
by: Li, Zhiyong, et al.
Published: (2026)
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
On the Reliability of Large Language Models for Causal Discovery
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Importance-Aware Data Augmentation for Document-Level Neural Machine Translation
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR
by: Wang, Minghan, et al.
Published: (2024)
by: Wang, Minghan, et al.
Published: (2024)
Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models
by: Wang, Minghan, et al.
Published: (2024)
by: Wang, Minghan, et al.
Published: (2024)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Similar Items
-
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024) -
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026) -
ReCA: A Parametric ReLU Composite Activation Function
by: Chidiac, John, et al.
Published: (2025) -
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
by: Wang, Weijie, et al.
Published: (2025) -
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)