MASIV: Toward Material-Agnostic System Identification from Videos
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Yizhou, Chen, Haoyu, Liu, Chunjiang, Li, Zhenyang, Herrmann, Charles, Hur, Junhwa, Li, Yinxiao, Yang, Ming-Hsuan, Raj, Bhiksha, Xu, Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MOSIV: Multi-Object System Identification from Videos
por: Liu, Chunjiang, et al.
Publicado: (2026)
por: Liu, Chunjiang, et al.
Publicado: (2026)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
por: Zhang, Junyi, et al.
Publicado: (2023)
por: Zhang, Junyi, et al.
Publicado: (2023)
Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting
por: Zhao, Yizhou, et al.
Publicado: (2025)
por: Zhao, Yizhou, et al.
Publicado: (2025)
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
por: Liu, Jinxiu, et al.
Publicado: (2024)
por: Liu, Jinxiu, et al.
Publicado: (2024)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
por: Zhang, Junyi, et al.
Publicado: (2024)
por: Zhang, Junyi, et al.
Publicado: (2024)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
por: Zhang, Junyi, et al.
Publicado: (2026)
por: Zhang, Junyi, et al.
Publicado: (2026)
Synergistic Global-space Camera and Human Reconstruction from Videos
por: Zhao, Yizhou, et al.
Publicado: (2024)
por: Zhao, Yizhou, et al.
Publicado: (2024)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
por: Gu, Leslie, et al.
Publicado: (2025)
por: Gu, Leslie, et al.
Publicado: (2025)
A Simple Approach to Unifying Diffusion-based Conditional Generation
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
por: Kim, Jihwan, et al.
Publicado: (2026)
por: Kim, Jihwan, et al.
Publicado: (2026)
Boundary Attention: Learning curves, corners, junctions and grouping
por: Polansky, Mia Gaia, et al.
Publicado: (2024)
por: Polansky, Mia Gaia, et al.
Publicado: (2024)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
por: Hur, Junhwa, et al.
Publicado: (2026)
por: Hur, Junhwa, et al.
Publicado: (2026)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
por: Wang, Xiaoyuan, et al.
Publicado: (2025)
por: Wang, Xiaoyuan, et al.
Publicado: (2025)
Human Voice is Unique
por: Singh, Rita, et al.
Publicado: (2025)
por: Singh, Rita, et al.
Publicado: (2025)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
por: Kim, Namho, et al.
Publicado: (2025)
por: Kim, Namho, et al.
Publicado: (2025)
Emergent Temporal Correspondences from Video Diffusion Transformers
por: Nam, Jisu, et al.
Publicado: (2025)
por: Nam, Jisu, et al.
Publicado: (2025)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
por: Nam, Jisu, et al.
Publicado: (2026)
por: Nam, Jisu, et al.
Publicado: (2026)
Lumiere: A Space-Time Diffusion Model for Video Generation
por: Bar-Tal, Omer, et al.
Publicado: (2024)
por: Bar-Tal, Omer, et al.
Publicado: (2024)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
por: Waheed, Abdul, et al.
Publicado: (2025)
por: Waheed, Abdul, et al.
Publicado: (2025)
Revisiting Acoustic Features for Robust ASR
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
por: Zhang, Hao, et al.
Publicado: (2023)
por: Zhang, Hao, et al.
Publicado: (2023)
QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition
por: Li, Xiang, et al.
Publicado: (2023)
por: Li, Xiang, et al.
Publicado: (2023)
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
por: Hur, Junhwa, et al.
Publicado: (2024)
por: Hur, Junhwa, et al.
Publicado: (2024)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
por: Chan, Kelvin C. K., et al.
Publicado: (2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories
por: Geng, Daniel, et al.
Publicado: (2024)
por: Geng, Daniel, et al.
Publicado: (2024)
Global Diffusive Expansion of Boltzmann Equation in exterior Domain
por: Jung, Junhwa
Publicado: (2023)
por: Jung, Junhwa
Publicado: (2023)
Evaluating and Improving Continual Learning in Spoken Language Understanding
por: Yang, Muqiao, et al.
Publicado: (2024)
por: Yang, Muqiao, et al.
Publicado: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
Domain Adaptation for Contrastive Audio-Language Models
por: Deshmukh, Soham, et al.
Publicado: (2024)
por: Deshmukh, Soham, et al.
Publicado: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
por: Cohen, Eyal, et al.
Publicado: (2025)
por: Cohen, Eyal, et al.
Publicado: (2025)
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
por: Baali, Massa, et al.
Publicado: (2025)
por: Baali, Massa, et al.
Publicado: (2025)
Repurposing Video Diffusion Transformers for Robust Point Tracking
por: Son, Soowon, et al.
Publicado: (2025)
por: Son, Soowon, et al.
Publicado: (2025)
Training-free CryoET Tomogram Segmentation
por: Zhao, Yizhou, et al.
Publicado: (2024)
por: Zhao, Yizhou, et al.
Publicado: (2024)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
por: Bao, Jingzhi, et al.
Publicado: (2024)
por: Bao, Jingzhi, et al.
Publicado: (2024)
GaMi: Geometry-Agnostic Material Identification via Cross-Modal Subtractive Disentanglement
por: Chen, Zhiwei, et al.
Publicado: (2026)
por: Chen, Zhiwei, et al.
Publicado: (2026)
WonderJourney: Going from Anywhere to Everywhere
por: Yu, Hong-Xing, et al.
Publicado: (2023)
por: Yu, Hong-Xing, et al.
Publicado: (2023)
Calibrated Multi-Preference Optimization for Aligning Diffusion Models
por: Lee, Kyungmin, et al.
Publicado: (2025)
por: Lee, Kyungmin, et al.
Publicado: (2025)
GR3EN: Generative Relighting for 3D Environments
por: Xing, Xiaoyan, et al.
Publicado: (2026)
por: Xing, Xiaoyan, et al.
Publicado: (2026)
Ejemplares similares
-
MOSIV: Multi-Object System Identification from Videos
por: Liu, Chunjiang, et al.
Publicado: (2026) -
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
por: Zhang, Junyi, et al.
Publicado: (2023) -
Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting
por: Zhao, Yizhou, et al.
Publicado: (2025) -
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
por: Liu, Jinxiu, et al.
Publicado: (2024) -
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
por: Zhang, Junyi, et al.
Publicado: (2024)