SIAM: A Simple Alternating Mixer for Video Prediction
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Xin, Peng, Ziang, Cao, Yuan, Shan, Hongming, Zhang, Junping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Shushing! Let's Imagine an Authentic Speech from the Silent Video
por: Ye, Jiaxin, et al.
Publicado: (2025)
por: Ye, Jiaxin, et al.
Publicado: (2025)
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
por: Zhao, Zilong, et al.
Publicado: (2026)
por: Zhao, Zilong, et al.
Publicado: (2026)
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
por: Huang, Zhongzhen, et al.
Publicado: (2024)
por: Huang, Zhongzhen, et al.
Publicado: (2024)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
por: Picard, David, et al.
Publicado: (2024)
por: Picard, David, et al.
Publicado: (2024)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
por: Yun, Seokju, et al.
Publicado: (2024)
por: Yun, Seokju, et al.
Publicado: (2024)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
por: Picard, David, et al.
Publicado: (2026)
por: Picard, David, et al.
Publicado: (2026)
KAN-Mixers: a new deep learning architecture for image classification
por: Canuto, Jorge Luiz dos Santos, et al.
Publicado: (2025)
por: Canuto, Jorge Luiz dos Santos, et al.
Publicado: (2025)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
por: Wang, Sibo, et al.
Publicado: (2024)
por: Wang, Sibo, et al.
Publicado: (2024)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
por: Niu, Muyao, et al.
Publicado: (2024)
por: Niu, Muyao, et al.
Publicado: (2024)
CausalVE: Face Video Privacy Encryption via Causal Video Prediction
por: Huang, Yubo, et al.
Publicado: (2024)
por: Huang, Yubo, et al.
Publicado: (2024)
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
por: Pan, Junwen, et al.
Publicado: (2025)
por: Pan, Junwen, et al.
Publicado: (2025)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
por: Zheng, Yikai, et al.
Publicado: (2026)
por: Zheng, Yikai, et al.
Publicado: (2026)
Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
por: Jiang, Changjiang, et al.
Publicado: (2025)
por: Jiang, Changjiang, et al.
Publicado: (2025)
Realism Control One-step Diffusion for Real-World Image Super-Resolution
por: Wu, Zongliang, et al.
Publicado: (2025)
por: Wu, Zongliang, et al.
Publicado: (2025)
TrajMamba: An Ego-Motion-Guided Mamba Model for Pedestrian Trajectory Prediction from an Egocentric Perspective
por: Peng, Yusheng, et al.
Publicado: (2026)
por: Peng, Yusheng, et al.
Publicado: (2026)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
por: Luo, Yuanhao, et al.
Publicado: (2026)
por: Luo, Yuanhao, et al.
Publicado: (2026)
Progressive Image Restoration via Text-Conditioned Video Generation
por: Kang, Peng, et al.
Publicado: (2025)
por: Kang, Peng, et al.
Publicado: (2025)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
por: Yang, Cheng, et al.
Publicado: (2025)
por: Yang, Cheng, et al.
Publicado: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
por: Bian, Yuxuan, et al.
Publicado: (2025)
por: Bian, Yuxuan, et al.
Publicado: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
por: Yu, Hu, et al.
Publicado: (2025)
por: Yu, Hu, et al.
Publicado: (2025)
DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
por: Ye, Jiaxin, et al.
Publicado: (2024)
por: Ye, Jiaxin, et al.
Publicado: (2024)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
por: Wu, Peng, et al.
Publicado: (2023)
por: Wu, Peng, et al.
Publicado: (2023)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
por: Yan, Bei, et al.
Publicado: (2024)
por: Yan, Bei, et al.
Publicado: (2024)
SeqTex: Generate Mesh Textures in Video Sequence
por: Yuan, Ze, et al.
Publicado: (2025)
por: Yuan, Ze, et al.
Publicado: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
por: Zhou, Mingde, et al.
Publicado: (2026)
por: Zhou, Mingde, et al.
Publicado: (2026)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
por: Chen, Zhihao, et al.
Publicado: (2023)
por: Chen, Zhihao, et al.
Publicado: (2023)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
por: Yang, Junqi, et al.
Publicado: (2026)
por: Yang, Junqi, et al.
Publicado: (2026)
Segment Anything for Videos: A Systematic Survey
por: Zhang, Chunhui, et al.
Publicado: (2024)
por: Zhang, Chunhui, et al.
Publicado: (2024)
UniSino: Physics-Driven Foundational Model for Universal CT Sinogram Standardization
por: Ai, Xingyu, et al.
Publicado: (2025)
por: Ai, Xingyu, et al.
Publicado: (2025)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
por: Li, Yuhang, et al.
Publicado: (2024)
por: Li, Yuhang, et al.
Publicado: (2024)
CutDiffusion: A Simple, Fast, Cheap, and Strong Diffusion Extrapolation Method
por: Lin, Mingbao, et al.
Publicado: (2024)
por: Lin, Mingbao, et al.
Publicado: (2024)
Deep Learning for Video Anomaly Detection: A Review
por: Wu, Peng, et al.
Publicado: (2024)
por: Wu, Peng, et al.
Publicado: (2024)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
por: Tian, Shulin, et al.
Publicado: (2025)
por: Tian, Shulin, et al.
Publicado: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
por: Fu, Hongming, et al.
Publicado: (2026)
por: Fu, Hongming, et al.
Publicado: (2026)
A Simple Aerial Detection Baseline of Multimodal Language Models
por: Li, Qingyun, et al.
Publicado: (2025)
por: Li, Qingyun, et al.
Publicado: (2025)
Point, Segment and Count: A Generalized Framework for Object Counting
por: Huang, Zhizhong, et al.
Publicado: (2023)
por: Huang, Zhizhong, et al.
Publicado: (2023)
VideoGuard: Protecting Video Content from Unauthorized Editing
por: Cao, Junjie, et al.
Publicado: (2025)
por: Cao, Junjie, et al.
Publicado: (2025)
BEVTrack: A Simple and Strong Baseline for 3D Single Object Tracking in Bird's-Eye View
por: Yang, Yuxiang, et al.
Publicado: (2023)
por: Yang, Yuxiang, et al.
Publicado: (2023)
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
por: Yuan, Liping, et al.
Publicado: (2025)
por: Yuan, Liping, et al.
Publicado: (2025)
Ejemplares similares
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video
por: Ye, Jiaxin, et al.
Publicado: (2025) -
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
por: Zhao, Zilong, et al.
Publicado: (2026) -
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
por: Huang, Zhongzhen, et al.
Publicado: (2024) -
PoM: Efficient Image and Video Generation with the Polynomial Mixer
por: Picard, David, et al.
Publicado: (2024) -
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
por: Yun, Seokju, et al.
Publicado: (2024)