Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Swain, Kabir, Han, Sijie, Weidele, Daniel Karl I., Martino, Mauro, Torralba, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
von: Zhang, Yuelin, et al.
Veröffentlicht: (2026)
von: Zhang, Yuelin, et al.
Veröffentlicht: (2026)
Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
von: Jiang, Wen, et al.
Veröffentlicht: (2025)
von: Jiang, Wen, et al.
Veröffentlicht: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
Automatic Discovery of Visual Circuits
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
ARM: Advantage Reward Modeling for Long-Horizon Manipulation
von: Mao, Yiming, et al.
Veröffentlicht: (2026)
von: Mao, Yiming, et al.
Veröffentlicht: (2026)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains
von: Zhan, Zhaohuan, et al.
Veröffentlicht: (2024)
von: Zhan, Zhaohuan, et al.
Veröffentlicht: (2024)
Pig aggression classification using CNN, Transformers and Recurrent Networks
von: Souza, Junior Silva, et al.
Veröffentlicht: (2024)
von: Souza, Junior Silva, et al.
Veröffentlicht: (2024)
RELIC: Interactive Video World Model with Long-Horizon Memory
von: Hong, Yicong, et al.
Veröffentlicht: (2025)
von: Hong, Yicong, et al.
Veröffentlicht: (2025)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
MultiModal Action Conditioned Video Generation
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models
von: Zhang, Ruizhi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruizhi, et al.
Veröffentlicht: (2026)
IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer
von: Qiu, Yuhang, et al.
Veröffentlicht: (2024)
von: Qiu, Yuhang, et al.
Veröffentlicht: (2024)
Block-Recurrent Dynamics in Vision Transformers
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
von: Kang, Inha, et al.
Veröffentlicht: (2025)
von: Kang, Inha, et al.
Veröffentlicht: (2025)
CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
von: Liu, Xuecong, et al.
Veröffentlicht: (2026)
von: Liu, Xuecong, et al.
Veröffentlicht: (2026)
bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition
von: Byra, Michal, et al.
Veröffentlicht: (2026)
von: Byra, Michal, et al.
Veröffentlicht: (2026)
POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion
von: Rigo, Andrea, et al.
Veröffentlicht: (2026)
von: Rigo, Andrea, et al.
Veröffentlicht: (2026)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
von: Wu, Ruiqi, et al.
Veröffentlicht: (2026)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2026)
MemoryMamba: Memory-Augmented State Space Model for Defect Recognition
von: Wang, Qianning, et al.
Veröffentlicht: (2024)
von: Wang, Qianning, et al.
Veröffentlicht: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
von: Chen, Haoyu, et al.
Veröffentlicht: (2026)
Learning Long-Term Temporal Dependencies in Photovoltaic Power Output Prediction Through Multi-Horizon Forecasting
von: Laha, Sumit, et al.
Veröffentlicht: (2026)
von: Laha, Sumit, et al.
Veröffentlicht: (2026)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
Single-pass Adaptive Image Tokenization for Minimum Program Search
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
von: Hosseintabar, Danial, et al.
Veröffentlicht: (2025)
von: Hosseintabar, Danial, et al.
Veröffentlicht: (2025)
Memory-Efficient Continual Learning Object Segmentation for Long Video
von: Nazemi, Amir, et al.
Veröffentlicht: (2023)
von: Nazemi, Amir, et al.
Veröffentlicht: (2023)
ReWind: Understanding Long Videos with Instructed Learnable Memory
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
von: Diko, Anxhelo, et al.
Veröffentlicht: (2024)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
VideoCAD: A Dataset and Model for Learning Long-Horizon 3D CAD UI Interactions from Video
von: Man, Brandon, et al.
Veröffentlicht: (2025)
von: Man, Brandon, et al.
Veröffentlicht: (2025)
Separating Knowledge and Perception with Procedural Data
von: Rodríguez-Muñoz, Adrián, et al.
Veröffentlicht: (2025)
von: Rodríguez-Muñoz, Adrián, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
von: Swain, Kabir, et al.
Veröffentlicht: (2026) -
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026) -
Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
von: Zhang, Yuelin, et al.
Veröffentlicht: (2026) -
Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation
von: Guo, Xinying, et al.
Veröffentlicht: (2026) -
Adaptive Length Image Tokenization via Recurrent Allocation
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)