Gespeichert in:
| Hauptverfasser: | Li, Longfei, Fan, Zhiwen, Cong, Wenyan, Liu, Xinhang, Yin, Yuyang, Foutter, Matt, Pan, Panwang, You, Chenyu, Wang, Yue, Wang, Zhangyang, Zhao, Yao, Pavone, Marco, Wei, Yunchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.07978 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
von: Yin, Yuyang, et al.
Veröffentlicht: (2023)
von: Yin, Yuyang, et al.
Veröffentlicht: (2023)
InstantSplat: Sparse-view Gaussian Splatting in Seconds
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
Can Test-Time Scaling Improve World Foundation Model?
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
von: Kwok, Jacky, et al.
Veröffentlicht: (2025)
von: Kwok, Jacky, et al.
Veröffentlicht: (2025)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
von: Xing, Ke, et al.
Veröffentlicht: (2025)
von: Xing, Ke, et al.
Veröffentlicht: (2025)
Real-Time Anomaly Detection and Reactive Planning with Large Language Models
von: Sinha, Rohan, et al.
Veröffentlicht: (2024)
von: Sinha, Rohan, et al.
Veröffentlicht: (2024)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
von: Liang, Hanwen, et al.
Veröffentlicht: (2024)
von: Liang, Hanwen, et al.
Veröffentlicht: (2024)
Learning Traffic Crashes as Language: Datasets, Benchmarks, and What-if Causal Analyses
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
von: Zhu, Zehao, et al.
Veröffentlicht: (2023)
von: Zhu, Zehao, et al.
Veröffentlicht: (2023)
Realistic Extreme Behavior Generation for Improved AV Testing
von: Dyro, Robert, et al.
Veröffentlicht: (2024)
von: Dyro, Robert, et al.
Veröffentlicht: (2024)
Vision Foundation Model Embedding-Based Semantic Anomaly Detection
von: Ronecker, Max Peter, et al.
Veröffentlicht: (2025)
von: Ronecker, Max Peter, et al.
Veröffentlicht: (2025)
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
von: Yin, Yuyang, et al.
Veröffentlicht: (2025)
von: Yin, Yuyang, et al.
Veröffentlicht: (2025)
4K4DGen: Panoramic 4D Generation at 4K Resolution
von: Li, Renjie, et al.
Veröffentlicht: (2024)
von: Li, Renjie, et al.
Veröffentlicht: (2024)
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
von: Wen, Kairun, et al.
Veröffentlicht: (2025)
von: Wen, Kairun, et al.
Veröffentlicht: (2025)
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
von: Foutter, Matthew, et al.
Veröffentlicht: (2024)
von: Foutter, Matthew, et al.
Veröffentlicht: (2024)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
von: Xiao, Yuxi, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxi, et al.
Veröffentlicht: (2025)
ReachBot Field Tests in a Mojave Desert Lava Tube as a Martian Analog
von: Chen, Tony G., et al.
Veröffentlicht: (2024)
von: Chen, Tony G., et al.
Veröffentlicht: (2024)
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
von: Li, Dayou, et al.
Veröffentlicht: (2026)
von: Li, Dayou, et al.
Veröffentlicht: (2026)
Martian Exploration of Lava Tubes (MELT) with ReachBot: Scientific Investigation and Concept of Operations
von: Di, Julia, et al.
Veröffentlicht: (2024)
von: Di, Julia, et al.
Veröffentlicht: (2024)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2023)
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2023)
CIPHER: Culvert Inspection through Pairwise Frame Selection and High-Efficiency Reconstruction
von: Lee, Seoyoung, et al.
Veröffentlicht: (2026)
von: Lee, Seoyoung, et al.
Veröffentlicht: (2026)
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
von: Yin, Li, et al.
Veröffentlicht: (2025)
von: Yin, Li, et al.
Veröffentlicht: (2025)
PlainQAFact: Retrieval-augmented Factual Consistency Evaluation Metric for Biomedical Plain Language Summarization
von: You, Zhiwen, et al.
Veröffentlicht: (2025)
von: You, Zhiwen, et al.
Veröffentlicht: (2025)
GaussianStego: A Generalizable Stenography Pipeline for Generative 3D Gaussians Splatting
von: Li, Chenxin, et al.
Veröffentlicht: (2024)
von: Li, Chenxin, et al.
Veröffentlicht: (2024)
CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy
von: Zhang, Jiakai, et al.
Veröffentlicht: (2025)
von: Zhang, Jiakai, et al.
Veröffentlicht: (2025)
Expressive Gaussian Human Avatars from Monocular RGB Video
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
Extrapolated Urban View Synthesis Benchmark
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior
von: Lin, Chenguo, et al.
Veröffentlicht: (2024)
von: Lin, Chenguo, et al.
Veröffentlicht: (2024)
PACE: Pacing Operator Learning to Accurate Optical Field Simulation for Complicated Photonic Devices
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models
von: Ren, Qin, et al.
Veröffentlicht: (2025)
von: Ren, Qin, et al.
Veröffentlicht: (2025)
A Stabilized High‐Order Spectral Model With Adaptive Residual‐Based Artificial Viscosity for Fully‐Nonlinear Free‐Surface Flow
von: Longfei Cong, et al.
Veröffentlicht: (2025)
von: Longfei Cong, et al.
Veröffentlicht: (2025)
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
InfoAffect: Affective Annotations of Infographics in Information Spread
von: Fu, Zihang, et al.
Veröffentlicht: (2025)
von: Fu, Zihang, et al.
Veröffentlicht: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2024)
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
von: Pan, Panwang, et al.
Veröffentlicht: (2025)
von: Pan, Panwang, et al.
Veröffentlicht: (2025)
Characterizing the current systems in the Martian ionosphere
von: Gao, Jiawei, et al.
Veröffentlicht: (2024)
von: Gao, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
von: Cong, Wenyan, et al.
Veröffentlicht: (2025) -
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
von: Yin, Yuyang, et al.
Veröffentlicht: (2023) -
InstantSplat: Sparse-view Gaussian Splatting in Seconds
von: Fan, Zhiwen, et al.
Veröffentlicht: (2024) -
Can Test-Time Scaling Improve World Foundation Model?
von: Cong, Wenyan, et al.
Veröffentlicht: (2025) -
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)