MSG Score: Automated Video Verification for Reliable Multi-Scene Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yoon, Daewon, Lee, Hyeongseok, Shin, Wonsik, Han, Sangyu, Kwak, Nojun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VDPP: Video Depth Post-Processing for Speed and Scalability
by: Yoon, Daewon, et al.
Published: (2026)
by: Yoon, Daewon, et al.
Published: (2026)
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
by: Han, Sangyu, et al.
Published: (2024)
by: Han, Sangyu, et al.
Published: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
by: Kim, Yearim, et al.
Published: (2024)
by: Kim, Yearim, et al.
Published: (2024)
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
by: Kim, Yearim, et al.
Published: (2026)
by: Kim, Yearim, et al.
Published: (2026)
S3D: Sketch-Driven 3D Model Generation
by: Song, Hail, et al.
Published: (2025)
by: Song, Hail, et al.
Published: (2025)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
by: Wang, Junxin, et al.
Published: (2026)
by: Wang, Junxin, et al.
Published: (2026)
LoGoColor: Local-Global 3D Colorization for 360° Scenes
by: Chang, Yeonjin, et al.
Published: (2025)
by: Chang, Yeonjin, et al.
Published: (2025)
Deep Polycuboid Fitting for Compact 3D Representation of Indoor Scenes
by: Lee, Gahye, et al.
Published: (2025)
by: Lee, Gahye, et al.
Published: (2025)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
by: Lin, Jiahao, et al.
Published: (2025)
by: Lin, Jiahao, et al.
Published: (2025)
Understanding Identity Continuity in Thermal Video through Scene-Level Consistency
by: Sun, Wei-Chieh, et al.
Published: (2026)
by: Sun, Wei-Chieh, et al.
Published: (2026)
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
by: Jang, Jae Won, et al.
Published: (2026)
by: Jang, Jae Won, et al.
Published: (2026)
Multi-resolution Rescored ByteTrack for Video Object Detection on Ultra-low-power Embedded Systems
by: Bompani, Luca, et al.
Published: (2024)
by: Bompani, Luca, et al.
Published: (2024)
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
by: Jiang, Zeren, et al.
Published: (2025)
by: Jiang, Zeren, et al.
Published: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
by: Jin, Haopeng, et al.
Published: (2026)
by: Jin, Haopeng, et al.
Published: (2026)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
by: Lee, Byung Hoon, et al.
Published: (2025)
by: Lee, Byung Hoon, et al.
Published: (2025)
SoccerNet-v3D: Leveraging Sports Broadcast Replays for 3D Scene Understanding
by: Gutiérrez-Pérez, Marc, et al.
Published: (2025)
by: Gutiérrez-Pérez, Marc, et al.
Published: (2025)
Detecting AI-Generated Videos with Spiking Neural Networks
by: Jang, Minsuk, et al.
Published: (2026)
by: Jang, Minsuk, et al.
Published: (2026)
RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety
by: Dumitriu, Andrei, et al.
Published: (2025)
by: Dumitriu, Andrei, et al.
Published: (2025)
Leveraging Out-of-Distribution Unlabeled Images: Semi-Supervised Semantic Segmentation with an Open-Vocabulary Model
by: Shin, Wooseok, et al.
Published: (2025)
by: Shin, Wooseok, et al.
Published: (2025)
Mono-Modalizing Extremely Heterogeneous Multi-Modal Medical Image Registration
by: Choo, Kyobin, et al.
Published: (2025)
by: Choo, Kyobin, et al.
Published: (2025)
Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization
by: Lee, Dongkwan, et al.
Published: (2025)
by: Lee, Dongkwan, et al.
Published: (2025)
Stochastic Video Generation with a Learned Prior
by: Denton, Remi, et al.
Published: (2018)
by: Denton, Remi, et al.
Published: (2018)
Efficient Verification-Based Face Identification
by: Rozner, Amit, et al.
Published: (2023)
by: Rozner, Amit, et al.
Published: (2023)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026)
by: Zheng, Hantao, et al.
Published: (2026)
Obtaining Favorable Layouts for Multiple Object Generation
by: Battash, Barak, et al.
Published: (2024)
by: Battash, Barak, et al.
Published: (2024)
ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos
by: Zhang, Zetong, et al.
Published: (2025)
by: Zhang, Zetong, et al.
Published: (2025)
Long Tail Image Generation Through Feature Space Augmentation and Iterated Learning
by: Elberg, Rafael, et al.
Published: (2024)
by: Elberg, Rafael, et al.
Published: (2024)
RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated Images
by: Yu, Hanzhe, et al.
Published: (2026)
by: Yu, Hanzhe, et al.
Published: (2026)
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models
by: Bastien, JF, et al.
Published: (2026)
by: Bastien, JF, et al.
Published: (2026)
Style Composition within Distinct LoRA modules for Traditional Art
by: Lee, Jaehyun, et al.
Published: (2025)
by: Lee, Jaehyun, et al.
Published: (2025)
CountPath: Automating Fragment Counting in Digital Pathology
by: Vieira, Ana Beatriz, et al.
Published: (2025)
by: Vieira, Ana Beatriz, et al.
Published: (2025)
M3LEO: A Multi-Modal, Multi-Label Earth Observation Dataset Integrating Interferometric SAR and Multispectral Data
by: Allen, Matthew J, et al.
Published: (2024)
by: Allen, Matthew J, et al.
Published: (2024)
Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
Two-step Authentication: Multi-biometric System Using Voice and Facial Recognition
by: Chen, Kuan Wei, et al.
Published: (2026)
by: Chen, Kuan Wei, et al.
Published: (2026)
EDSNet: Efficient-DSNet for Video Summarization
by: Prasad, Ashish, et al.
Published: (2024)
by: Prasad, Ashish, et al.
Published: (2024)
Intrinsic Image Fusion for Multi-View 3D Material Reconstruction
by: Kocsis, Peter, et al.
Published: (2025)
by: Kocsis, Peter, et al.
Published: (2025)
Content Adaptive based Motion Alignment Framework for Learned Video Compression
by: Zhang, Tiange, et al.
Published: (2025)
by: Zhang, Tiange, et al.
Published: (2025)
A unified Benchmark for Multi-Frame Image Restoration under Severe Refractive Warping
by: Shugaev, Maxim V., et al.
Published: (2026)
by: Shugaev, Maxim V., et al.
Published: (2026)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
by: Dokme, Atahan, et al.
Published: (2026)
by: Dokme, Atahan, et al.
Published: (2026)
Similar Items
-
VDPP: Video Depth Post-Processing for Speed and Scalability
by: Yoon, Daewon, et al.
Published: (2026) -
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
by: Han, Sangyu, et al.
Published: (2024) -
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025) -
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
by: Kim, Yearim, et al.
Published: (2024) -
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
by: Kim, Yearim, et al.
Published: (2026)