Saved in:
| Main Authors: | Zheng, Xin, Peng, Ziang, Cao, Yuan, Shan, Hongming, Zhang, Junping |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.11683 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shushing! Let's Imagine an Authentic Speech from the Silent Video
by: Ye, Jiaxin, et al.
Published: (2025)
by: Ye, Jiaxin, et al.
Published: (2025)
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
by: Zhao, Zilong, et al.
Published: (2026)
by: Zhao, Zilong, et al.
Published: (2026)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024)
by: Picard, David, et al.
Published: (2024)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
by: Picard, David, et al.
Published: (2026)
by: Picard, David, et al.
Published: (2026)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
by: Ye, Jiaxin, et al.
Published: (2024)
by: Ye, Jiaxin, et al.
Published: (2024)
KAN-Mixers: a new deep learning architecture for image classification
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
by: Zheng, Yikai, et al.
Published: (2026)
by: Zheng, Yikai, et al.
Published: (2026)
CausalVE: Face Video Privacy Encryption via Causal Video Prediction
by: Huang, Yubo, et al.
Published: (2024)
by: Huang, Yubo, et al.
Published: (2024)
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
by: Jiang, Changjiang, et al.
Published: (2025)
by: Jiang, Changjiang, et al.
Published: (2025)
Realism Control One-step Diffusion for Real-World Image Super-Resolution
by: Wu, Zongliang, et al.
Published: (2025)
by: Wu, Zongliang, et al.
Published: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
Point, Segment and Count: A Generalized Framework for Object Counting
by: Huang, Zhizhong, et al.
Published: (2023)
by: Huang, Zhizhong, et al.
Published: (2023)
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
by: Chen, Zhihao, et al.
Published: (2023)
by: Chen, Zhihao, et al.
Published: (2023)
UniSino: Physics-Driven Foundational Model for Universal CT Sinogram Standardization
by: Ai, Xingyu, et al.
Published: (2025)
by: Ai, Xingyu, et al.
Published: (2025)
SeqTex: Generate Mesh Textures in Video Sequence
by: Yuan, Ze, et al.
Published: (2025)
by: Yuan, Ze, et al.
Published: (2025)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
TrajMamba: An Ego-Motion-Guided Mamba Model for Pedestrian Trajectory Prediction from an Egocentric Perspective
by: Peng, Yusheng, et al.
Published: (2026)
by: Peng, Yusheng, et al.
Published: (2026)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
by: Yan, Bei, et al.
Published: (2024)
by: Yan, Bei, et al.
Published: (2024)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
MVR: Multi-view Video Reward Shaping for Reinforcement Learning
by: Luo, Lirui, et al.
Published: (2026)
by: Luo, Lirui, et al.
Published: (2026)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)
by: Tian, Shulin, et al.
Published: (2025)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
by: Yang, Junqi, et al.
Published: (2026)
by: Yang, Junqi, et al.
Published: (2026)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
CutDiffusion: A Simple, Fast, Cheap, and Strong Diffusion Extrapolation Method
by: Lin, Mingbao, et al.
Published: (2024)
by: Lin, Mingbao, et al.
Published: (2024)
Research on Intelligent Aided Diagnosis System of Medical Image Based on Computer Deep Learning
by: Yuan, Jiajie, et al.
Published: (2024)
by: Yuan, Jiajie, et al.
Published: (2024)
A Simple Aerial Detection Baseline of Multimodal Language Models
by: Li, Qingyun, et al.
Published: (2025)
by: Li, Qingyun, et al.
Published: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
by: Zhou, Mingde, et al.
Published: (2026)
by: Zhou, Mingde, et al.
Published: (2026)
Deep Learning for Video Anomaly Detection: A Review
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
BEVTrack: A Simple and Strong Baseline for 3D Single Object Tracking in Bird's-Eye View
by: Yang, Yuxiang, et al.
Published: (2023)
by: Yang, Yuxiang, et al.
Published: (2023)
Similar Items
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video
by: Ye, Jiaxin, et al.
Published: (2025) -
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
by: Zhao, Zilong, et al.
Published: (2026) -
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024) -
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
by: Yun, Seokju, et al.
Published: (2024) -
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)