Gespeichert in:
| Hauptverfasser: | Huang, Zixuan, Li, Xiang, Lv, Zhaoyang, Rehg, James M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2512.19949 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vinedresser3D: Agentic Text-guided 3D Editing
von: Chi, Yankuan, et al.
Veröffentlicht: (2026)
von: Chi, Yankuan, et al.
Veröffentlicht: (2026)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
von: Li, Kunhang, et al.
Veröffentlicht: (2025)
von: Li, Kunhang, et al.
Veröffentlicht: (2025)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
von: Kim, Junho, et al.
Veröffentlicht: (2026)
von: Kim, Junho, et al.
Veröffentlicht: (2026)
MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
von: Tong, Jinguang, et al.
Veröffentlicht: (2026)
von: Tong, Jinguang, et al.
Veröffentlicht: (2026)
Yan: Foundational Interactive Video Generation
von: Ye, Deheng, et al.
Veröffentlicht: (2025)
von: Ye, Deheng, et al.
Veröffentlicht: (2025)
How Much You Ate? Food Portion Estimation on Spoons
von: Sharma, Aaryam, et al.
Veröffentlicht: (2024)
von: Sharma, Aaryam, et al.
Veröffentlicht: (2024)
Guess the Unified Model: How Much Can We Recover from Generated Images?
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2026)
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2026)
Medical Video Generation for Disease Progression Simulation
von: Cao, Xu, et al.
Veröffentlicht: (2024)
von: Cao, Xu, et al.
Veröffentlicht: (2024)
SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation
von: Tang, Zixuan, et al.
Veröffentlicht: (2026)
von: Tang, Zixuan, et al.
Veröffentlicht: (2026)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
Do Pre-trained Vision-Language Models Encode Object States?
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
von: Newman, Kaleb, et al.
Veröffentlicht: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
von: Gao, Zihui, et al.
Veröffentlicht: (2026)
von: Gao, Zihui, et al.
Veröffentlicht: (2026)
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
von: Huang, Zixuan, et al.
Veröffentlicht: (2025)
von: Huang, Zixuan, et al.
Veröffentlicht: (2025)
Does Semantic Noise Initialization Transfer from Images to Videos? A Paired Diagnostic Study
von: Jing, Yixiao, et al.
Veröffentlicht: (2026)
von: Jing, Yixiao, et al.
Veröffentlicht: (2026)
Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
von: Adorni, Pierre, et al.
Veröffentlicht: (2025)
von: Adorni, Pierre, et al.
Veröffentlicht: (2025)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
A Generative Foundation Model for Multimodal Histopathology
von: Xiang, Jinxi, et al.
Veröffentlicht: (2026)
von: Xiang, Jinxi, et al.
Veröffentlicht: (2026)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jihyeon, et al.
Veröffentlicht: (2026)
MetaSSC: Enhancing 3D Semantic Scene Completion for Autonomous Driving through Meta-Learning and Long-sequence Modeling
von: Qu, Yansong, et al.
Veröffentlicht: (2024)
von: Qu, Yansong, et al.
Veröffentlicht: (2024)
FMGS: Foundation Model Embedded 3D Gaussian Splatting for Holistic 3D Scene Understanding
von: Zuo, Xingxing, et al.
Veröffentlicht: (2024)
von: Zuo, Xingxing, et al.
Veröffentlicht: (2024)
How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
von: Chang, Chirui, et al.
Veröffentlicht: (2024)
von: Chang, Chirui, et al.
Veröffentlicht: (2024)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
von: Zheng, Duo, et al.
Veröffentlicht: (2025)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
von: Huang, YiJie, et al.
Veröffentlicht: (2026)
von: Huang, YiJie, et al.
Veröffentlicht: (2026)
Triad: Vision Foundation Model for 3D Magnetic Resonance Imaging
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters
von: Yu, Yinghong, et al.
Veröffentlicht: (2026)
von: Yu, Yinghong, et al.
Veröffentlicht: (2026)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
von: Linghu, Xiongkun, et al.
Veröffentlicht: (2026)
SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input
von: Lv, Zhen, et al.
Veröffentlicht: (2024)
von: Lv, Zhen, et al.
Veröffentlicht: (2024)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography
von: Zhu, Weicheng, et al.
Veröffentlicht: (2025)
von: Zhu, Weicheng, et al.
Veröffentlicht: (2025)
M$^3$-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
Few-shot Semantic Encoding and Decoding for Video Surveillance
von: Cheng, Baoping, et al.
Veröffentlicht: (2025)
von: Cheng, Baoping, et al.
Veröffentlicht: (2025)
BehAVE: Behaviour Alignment of Video Game Encodings
von: Rašajski, Nemanja, et al.
Veröffentlicht: (2024)
von: Rašajski, Nemanja, et al.
Veröffentlicht: (2024)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
von: Li, Lingen, et al.
Veröffentlicht: (2026)
von: Li, Lingen, et al.
Veröffentlicht: (2026)
LoV3D: Grounding Cognitive Prognosis Reasoning in Longitudinal 3D Brain MRI via Regional Volume Assessments
von: Jiang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Jiang, Zhaoyang, et al.
Veröffentlicht: (2026)
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
von: Prakash, Aditya, et al.
Veröffentlicht: (2025)
von: Prakash, Aditya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Vinedresser3D: Agentic Text-guided 3D Editing
von: Chi, Yankuan, et al.
Veröffentlicht: (2026) -
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2025) -
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2024) -
How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
von: Li, Kunhang, et al.
Veröffentlicht: (2025) -
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
von: Kim, Junho, et al.
Veröffentlicht: (2026)