Saved in:
| Main Authors: | Sun, Yang-Che, Sun, Cheng, Lin, Chin-Yang, Yang, Fu-En, Chen, Min-Hung, Lin, Yen-Yu, Liu, Yu-Lun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.08831 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
by: Lin, Chin-Yang, et al.
Published: (2025)
by: Lin, Chin-Yang, et al.
Published: (2025)
CorrFill: Enhancing Faithfulness in Reference-based Inpainting with Correspondence Guidance in Diffusion Models
by: Liu, Kuan-Hung, et al.
Published: (2025)
by: Liu, Kuan-Hung, et al.
Published: (2025)
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
by: Cheng, Ying, et al.
Published: (2025)
by: Cheng, Ying, et al.
Published: (2025)
Segment Anything, Even Occluded
by: Tai, Wei-En, et al.
Published: (2025)
by: Tai, Wei-En, et al.
Published: (2025)
FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors
by: Lin, Chin-Yang, et al.
Published: (2024)
by: Lin, Chin-Yang, et al.
Published: (2024)
2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
by: Yang, Cheng-Kun, et al.
Published: (2023)
by: Yang, Cheng-Kun, et al.
Published: (2023)
PartDistill: 3D Shape Part Segmentation by Vision-Language Model Distillation
by: Umam, Ardian, et al.
Published: (2023)
by: Umam, Ardian, et al.
Published: (2023)
HDR Reconstruction Boosting with Training-Free and Exposure-Consistent Diffusion
by: Lin, Yo-Tin, et al.
Published: (2026)
by: Lin, Yo-Tin, et al.
Published: (2026)
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction
by: Huang, Yi-Chuan, et al.
Published: (2025)
by: Huang, Yi-Chuan, et al.
Published: (2025)
FIPER: Factorized Features for Robust Image Super-Resolution and Compression
by: Sun, Yang-Che, et al.
Published: (2024)
by: Sun, Yang-Che, et al.
Published: (2024)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting
by: Wu, Chung-Ho, et al.
Published: (2025)
by: Wu, Chung-Ho, et al.
Published: (2025)
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
by: Shiu, Hau-Shiang, et al.
Published: (2025)
by: Shiu, Hau-Shiang, et al.
Published: (2025)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
by: Chiang, Jui-Che, et al.
Published: (2024)
by: Chiang, Jui-Che, et al.
Published: (2024)
OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
by: Hsu, Peng-Hao, et al.
Published: (2025)
by: Hsu, Peng-Hao, et al.
Published: (2025)
Measuring 3D Spatial Geometric Consistency in Dynamic Generated Videos
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation
by: Lee, Han-Hung, et al.
Published: (2026)
by: Lee, Han-Hung, et al.
Published: (2026)
Place Anything into Any Video
by: Liu, Ziling, et al.
Published: (2024)
by: Liu, Ziling, et al.
Published: (2024)
DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models
by: Yeh, Chang-Han, et al.
Published: (2024)
by: Yeh, Chang-Han, et al.
Published: (2024)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
by: Ma, Junxian, et al.
Published: (2025)
by: Ma, Junxian, et al.
Published: (2025)
Deep Geometric Moments Promote Shape Consistency in Text-to-3D Generation
by: Nath, Utkarsh, et al.
Published: (2024)
by: Nath, Utkarsh, et al.
Published: (2024)
DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Detect Anything 3D in the Wild
by: Zhang, Hanxue, et al.
Published: (2025)
by: Zhang, Hanxue, et al.
Published: (2025)
GenRC: Generative 3D Room Completion from Sparse Image Collections
by: Li, Ming-Feng, et al.
Published: (2024)
by: Li, Ming-Feng, et al.
Published: (2024)
ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
by: Chen, Jr-Jen, et al.
Published: (2024)
by: Chen, Jr-Jen, et al.
Published: (2024)
GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth
by: Liu, Yuecheng, et al.
Published: (2026)
by: Liu, Yuecheng, et al.
Published: (2026)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing
by: Liu, Feng-Lin, et al.
Published: (2025)
by: Liu, Feng-Lin, et al.
Published: (2025)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
by: Sun, Yang-Tian, et al.
Published: (2025)
by: Sun, Yang-Tian, et al.
Published: (2025)
RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes
by: Lee, Yuan-Kang, et al.
Published: (2026)
by: Lee, Yuan-Kang, et al.
Published: (2026)
LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
TS-SAM: Fine-Tuning Segment-Anything Model for Downstream Tasks
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
AnimateAnything: Consistent and Controllable Animation for Video Generation
by: Lei, Guojun, et al.
Published: (2024)
by: Lei, Guojun, et al.
Published: (2024)
LoBE-GS: Load-Balanced and Efficient 3D Gaussian Splatting for Large-Scale Scene Reconstruction
by: Hung, Sheng-Hsiang, et al.
Published: (2025)
by: Hung, Sheng-Hsiang, et al.
Published: (2025)
V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations
by: Jhang, Jin-Cheng, et al.
Published: (2024)
by: Jhang, Jin-Cheng, et al.
Published: (2024)
Similar Items
-
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
by: Lin, Chin-Yang, et al.
Published: (2025) -
CorrFill: Enhancing Faithfulness in Reference-based Inpainting with Correspondence Guidance in Diffusion Models
by: Liu, Kuan-Hung, et al.
Published: (2025) -
VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
by: Cheng, Ying, et al.
Published: (2025) -
Segment Anything, Even Occluded
by: Tai, Wei-En, et al.
Published: (2025) -
FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors
by: Lin, Chin-Yang, et al.
Published: (2024)