Saved in:
| Main Authors: | Huang, Zixuan, Li, Xiang, Lv, Zhaoyang, Rehg, James M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.19949 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vinedresser3D: Agentic Text-guided 3D Editing
by: Chi, Yankuan, et al.
Published: (2026)
by: Chi, Yankuan, et al.
Published: (2026)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
by: Li, Kunhang, et al.
Published: (2025)
by: Li, Kunhang, et al.
Published: (2025)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
by: Tong, Jinguang, et al.
Published: (2026)
by: Tong, Jinguang, et al.
Published: (2026)
Yan: Foundational Interactive Video Generation
by: Ye, Deheng, et al.
Published: (2025)
by: Ye, Deheng, et al.
Published: (2025)
How Much You Ate? Food Portion Estimation on Spoons
by: Sharma, Aaryam, et al.
Published: (2024)
by: Sharma, Aaryam, et al.
Published: (2024)
Guess the Unified Model: How Much Can We Recover from Generated Images?
by: Cekinmez, Jasin, et al.
Published: (2026)
by: Cekinmez, Jasin, et al.
Published: (2026)
Medical Video Generation for Disease Progression Simulation
by: Cao, Xu, et al.
Published: (2024)
by: Cao, Xu, et al.
Published: (2024)
SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation
by: Tang, Zixuan, et al.
Published: (2026)
by: Tang, Zixuan, et al.
Published: (2026)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
by: Yuan, Chao, et al.
Published: (2026)
by: Yuan, Chao, et al.
Published: (2026)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
by: Gao, Zihui, et al.
Published: (2026)
by: Gao, Zihui, et al.
Published: (2026)
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Does Semantic Noise Initialization Transfer from Images to Videos? A Paired Diagnostic Study
by: Jing, Yixiao, et al.
Published: (2026)
by: Jing, Yixiao, et al.
Published: (2026)
Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
by: Adorni, Pierre, et al.
Published: (2025)
by: Adorni, Pierre, et al.
Published: (2025)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
A Generative Foundation Model for Multimodal Histopathology
by: Xiang, Jinxi, et al.
Published: (2026)
by: Xiang, Jinxi, et al.
Published: (2026)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
by: Kim, Jihyeon, et al.
Published: (2026)
by: Kim, Jihyeon, et al.
Published: (2026)
MetaSSC: Enhancing 3D Semantic Scene Completion for Autonomous Driving through Meta-Learning and Long-sequence Modeling
by: Qu, Yansong, et al.
Published: (2024)
by: Qu, Yansong, et al.
Published: (2024)
FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
by: Zuo, Xingxing, et al.
Published: (2024)
by: Zuo, Xingxing, et al.
Published: (2024)
How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
by: Chang, Chirui, et al.
Published: (2024)
by: Chang, Chirui, et al.
Published: (2024)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
by: Yu, Zhuang, et al.
Published: (2026)
by: Yu, Zhuang, et al.
Published: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
by: Huang, YiJie, et al.
Published: (2026)
by: Huang, YiJie, et al.
Published: (2026)
Triad: Vision Foundation Model for 3D Magnetic Resonance Imaging
by: Wang, Shansong, et al.
Published: (2025)
by: Wang, Shansong, et al.
Published: (2025)
PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters
by: Yu, Yinghong, et al.
Published: (2026)
by: Yu, Yinghong, et al.
Published: (2026)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
by: Linghu, Xiongkun, et al.
Published: (2026)
by: Linghu, Xiongkun, et al.
Published: (2026)
SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input
by: Lv, Zhen, et al.
Published: (2024)
by: Lv, Zhen, et al.
Published: (2024)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026)
by: Fan, Yingqi, et al.
Published: (2026)
3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography
by: Zhu, Weicheng, et al.
Published: (2025)
by: Zhu, Weicheng, et al.
Published: (2025)
M$^3$-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
Few-shot Semantic Encoding and Decoding for Video Surveillance
by: Cheng, Baoping, et al.
Published: (2025)
by: Cheng, Baoping, et al.
Published: (2025)
BehAVE: Behaviour Alignment of Video Game Encodings
by: Rašajski, Nemanja, et al.
Published: (2024)
by: Rašajski, Nemanja, et al.
Published: (2024)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
by: Li, Lingen, et al.
Published: (2026)
by: Li, Lingen, et al.
Published: (2026)
LoV3D: Grounding Cognitive Prognosis Reasoning in Longitudinal 3D Brain MRI via Regional Volume Assessments
by: Jiang, Zhaoyang, et al.
Published: (2026)
by: Jiang, Zhaoyang, et al.
Published: (2026)
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
by: Prakash, Aditya, et al.
Published: (2025)
by: Prakash, Aditya, et al.
Published: (2025)
Similar Items
-
Vinedresser3D: Agentic Text-guided 3D Editing
by: Chi, Yankuan, et al.
Published: (2026) -
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
by: Li, Xiang, et al.
Published: (2025) -
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
by: Li, Xiang, et al.
Published: (2024) -
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
by: Li, Kunhang, et al.
Published: (2025) -
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)