MEMFOF: High-Resolution Training for Memory-Efficient Multi-Frame Optical Flow Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bargatin, Vladislav, Chistov, Egor, Yakovenko, Alexander, Vatolin, Dmitriy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
Color Mismatches in Stereoscopic Video: Real-World Dataset and Deep Correction Method
von: Chistov, Egor, et al.
Veröffentlicht: (2023)
von: Chistov, Egor, et al.
Veröffentlicht: (2023)
Can No-Reference Quality-Assessment Methods Serve as Perceptual Losses for Super-Resolution?
von: Kashkarov, Egor, et al.
Veröffentlicht: (2024)
von: Kashkarov, Egor, et al.
Veröffentlicht: (2024)
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Rethink Predicting the Optical Flow with the Kinetics Perspective
von: Cheng, Yuhao, et al.
Veröffentlicht: (2024)
von: Cheng, Yuhao, et al.
Veröffentlicht: (2024)
Efficient Low-Resolution Face Recognition via Bridge Distillation
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
von: Dang, Yunkai, et al.
Veröffentlicht: (2025)
von: Dang, Yunkai, et al.
Veröffentlicht: (2025)
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
von: Li, Shenshen, et al.
Veröffentlicht: (2025)
von: Li, Shenshen, et al.
Veröffentlicht: (2025)
NIC-RobustBench: A Comprehensive Open-Source Toolkit for Neural Image Compression and Robustness Analysis
von: Bychkov, Georgii, et al.
Veröffentlicht: (2025)
von: Bychkov, Georgii, et al.
Veröffentlicht: (2025)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2023)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2023)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
von: Zhang, Junzheng, et al.
Veröffentlicht: (2024)
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Text-Only Data Synthesis for Vision Language Model Training
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
Harmonizing Attention: Training-free Texture-aware Geometry Transfer
von: Ikuta, Eito, et al.
Veröffentlicht: (2024)
von: Ikuta, Eito, et al.
Veröffentlicht: (2024)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
von: Liu, Che, et al.
Veröffentlicht: (2026)
von: Liu, Che, et al.
Veröffentlicht: (2026)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
Video Seal: Open and Efficient Video Watermarking
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2024)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
von: Li, Liupeng, et al.
Veröffentlicht: (2026)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
von: Liu, Ke, et al.
Veröffentlicht: (2026)
von: Liu, Ke, et al.
Veröffentlicht: (2026)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Hai-Tao, et al.
Veröffentlicht: (2024)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
von: Zhang, Kangkai, et al.
Veröffentlicht: (2024)
von: Zhang, Kangkai, et al.
Veröffentlicht: (2024)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
von: Moser, Brian B., et al.
Veröffentlicht: (2024)
von: Moser, Brian B., et al.
Veröffentlicht: (2024)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
von: Xia, Yichong, et al.
Veröffentlicht: (2026)
von: Xia, Yichong, et al.
Veröffentlicht: (2026)
Robust Fuzzy Multi-view Learning under View Conflict
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
von: Tsereh, Daria, et al.
Veröffentlicht: (2024) -
Color Mismatches in Stereoscopic Video: Real-World Dataset and Deep Correction Method
von: Chistov, Egor, et al.
Veröffentlicht: (2023) -
Can No-Reference Quality-Assessment Methods Serve as Perceptual Losses for Super-Resolution?
von: Kashkarov, Egor, et al.
Veröffentlicht: (2024) -
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2025) -
Rethink Predicting the Optical Flow with the Kinetics Perspective
von: Cheng, Yuhao, et al.
Veröffentlicht: (2024)