Saved in:
| Main Authors: | Lee, Junho, Shin, Jeongwoo, Ko, Seung Woo, Ha, Seongsu, Lee, Joonseok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.05260 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025)
by: Shin, Jeongwoo, et al.
Published: (2025)
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
by: Lee, Junho, et al.
Published: (2025)
by: Lee, Junho, et al.
Published: (2025)
Geometry-Aware Image Flow Matching
by: Lee, Junho, et al.
Published: (2026)
by: Lee, Junho, et al.
Published: (2026)
Isometric Representation Learning for Disentangled Latent Space of Diffusion Models
by: Hahm, Jaehoon, et al.
Published: (2024)
by: Hahm, Jaehoon, et al.
Published: (2024)
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
by: Shin, Jeongwoo, et al.
Published: (2026)
by: Shin, Jeongwoo, et al.
Published: (2026)
GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
by: Lee, Inseo, et al.
Published: (2025)
by: Lee, Inseo, et al.
Published: (2025)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
by: Ryu, Suho, et al.
Published: (2025)
by: Ryu, Suho, et al.
Published: (2025)
A Revisit to the Decoder for Camouflaged Object Detection
by: Ko, Seung Woo, et al.
Published: (2025)
by: Ko, Seung Woo, et al.
Published: (2025)
Modality-Aware Representation Learning for Zero-shot Sketch-based Image Retrieval
by: Lyou, Eunyi, et al.
Published: (2024)
by: Lyou, Eunyi, et al.
Published: (2024)
ProDepth: Boosting Self-Supervised Multi-Frame Monocular Depth with Probabilistic Fusion
by: Woo, Sungmin, et al.
Published: (2024)
by: Woo, Sungmin, et al.
Published: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
by: Ahn, Geo, et al.
Published: (2026)
by: Ahn, Geo, et al.
Published: (2026)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
by: Lee, Hosu, et al.
Published: (2024)
by: Lee, Hosu, et al.
Published: (2024)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
by: Hwang, Geunmin, et al.
Published: (2025)
by: Hwang, Geunmin, et al.
Published: (2025)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
by: Lee, Hosu, et al.
Published: (2025)
by: Lee, Hosu, et al.
Published: (2025)
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
by: Pan, Yulu, et al.
Published: (2026)
by: Pan, Yulu, et al.
Published: (2026)
Fitting Image Diffusion Models on Video Datasets
by: Lee, Juhun, et al.
Published: (2025)
by: Lee, Juhun, et al.
Published: (2025)
An Adaptive Method Stabilizing Activations for Enhanced Generalization
by: Seung, Hyunseok, et al.
Published: (2025)
by: Seung, Hyunseok, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
H2G: Hierarchy-Aware Hyperbolic Grouping for 3D Scenes
by: Ko, ByungHa, et al.
Published: (2026)
by: Ko, ByungHa, et al.
Published: (2026)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
CAT: Class Aware Adaptive Thresholding for Semi-Supervised Domain Generalization
by: Zoha, Sumaiya, et al.
Published: (2024)
by: Zoha, Sumaiya, et al.
Published: (2024)
Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
by: Jeon, Wooseok, et al.
Published: (2026)
by: Jeon, Wooseok, et al.
Published: (2026)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
by: Kim, Sumin, et al.
Published: (2026)
by: Kim, Sumin, et al.
Published: (2026)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
by: Lee, Jeongwoo, et al.
Published: (2024)
by: Lee, Jeongwoo, et al.
Published: (2024)
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
Shot-Aware Frame Sampling for Video Understanding
by: Zhao, Mengyu, et al.
Published: (2026)
by: Zhao, Mengyu, et al.
Published: (2026)
Video Diffusion Models are Strong Video Inpainter
by: Lee, Minhyeok, et al.
Published: (2024)
by: Lee, Minhyeok, et al.
Published: (2024)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
by: Lee, Eunsoo, et al.
Published: (2026)
by: Lee, Eunsoo, et al.
Published: (2026)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Curriculum Fine-tuning of Vision Foundation Model for Medical Image Classification Under Label Noise
by: Yu, Yeonguk, et al.
Published: (2024)
by: Yu, Yeonguk, et al.
Published: (2024)
AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
by: Kye, Dahyeon, et al.
Published: (2025)
by: Kye, Dahyeon, et al.
Published: (2025)
Harnessing Meta-Learning for Improving Full-Frame Video Stabilization
by: Ali, Muhammad Kashif, et al.
Published: (2024)
by: Ali, Muhammad Kashif, et al.
Published: (2024)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
CAST: Cross-Attention in Space and Time for Video Action Recognition
by: Lee, Dongho, et al.
Published: (2023)
by: Lee, Dongho, et al.
Published: (2023)
Geometric Remove-and-Retrain (GOAR): Coordinate-Invariant eXplainable AI Assessment
by: Park, Yong-Hyun, et al.
Published: (2024)
by: Park, Yong-Hyun, et al.
Published: (2024)
Similar Items
-
Self-Guided Masked Autoencoder
by: Shin, Jeongwoo, et al.
Published: (2025) -
Latent Diffusion Models with Masked AutoEncoders
by: Lee, Junho, et al.
Published: (2025) -
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024) -
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025) -
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)