Gespeichert in:
| Hauptverfasser: | Shen, Hanwen, Lu, Jiajie, Cao, Yupeng, Yang, Xiaonan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.18046 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
Diffusion Adversarial Post-Training for One-Step Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
von: Zhang, Bowen, et al.
Veröffentlicht: (2024)
von: Zhang, Bowen, et al.
Veröffentlicht: (2024)
MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
Post-Training Quantization for Video Matting
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training
von: Liu, Anglin, et al.
Veröffentlicht: (2026)
von: Liu, Anglin, et al.
Veröffentlicht: (2026)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
von: Yu, An, et al.
Veröffentlicht: (2025)
von: Yu, An, et al.
Veröffentlicht: (2025)
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation
von: Chen, Zini, et al.
Veröffentlicht: (2026)
von: Chen, Zini, et al.
Veröffentlicht: (2026)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
von: Le, Huy, et al.
Veröffentlicht: (2025)
von: Le, Huy, et al.
Veröffentlicht: (2025)
OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation
von: Wang, Cong, et al.
Veröffentlicht: (2024)
von: Wang, Cong, et al.
Veröffentlicht: (2024)
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
von: Lu, Yifan, et al.
Veröffentlicht: (2024)
von: Lu, Yifan, et al.
Veröffentlicht: (2024)
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
von: Kang, Xueyang, et al.
Veröffentlicht: (2025)
von: Kang, Xueyang, et al.
Veröffentlicht: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
D3: Training-Free AI-Generated Video Detection Using Second-Order Features
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
von: Zheng, Chende, et al.
Veröffentlicht: (2025)
Enhancing Video Summarization with Context Awareness
von: Huynh-Lam, Hai-Dang, et al.
Veröffentlicht: (2024)
von: Huynh-Lam, Hai-Dang, et al.
Veröffentlicht: (2024)
DynamicPAE: Generating Scene-Aware Physical Adversarial Examples in Real-Time
von: Hu, Jin, et al.
Veröffentlicht: (2024)
von: Hu, Jin, et al.
Veröffentlicht: (2024)
PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2023)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2023)
TSTMotion: Training-free Scene-aware Text-to-motion Generation
von: Guo, Ziyan, et al.
Veröffentlicht: (2025)
von: Guo, Ziyan, et al.
Veröffentlicht: (2025)
Training-free Composite Scene Generation for Layout-to-Image Synthesis
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
Pre-Trained Video Generative Models as World Simulators
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions
von: Yu, Jiashuo, et al.
Veröffentlicht: (2026)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2026)
Video-Bench: Human-Aligned Video Generation Benchmark
von: Han, Hui, et al.
Veröffentlicht: (2025)
von: Han, Hui, et al.
Veröffentlicht: (2025)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
von: Zhang, Shengjun, et al.
Veröffentlicht: (2025)
von: Zhang, Shengjun, et al.
Veröffentlicht: (2025)
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
von: Li, Jiajie, et al.
Veröffentlicht: (2024)
Revisit Human-Scene Interaction via Space Occupancy
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
von: Liu, Xinpeng, et al.
Veröffentlicht: (2023)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
von: Shen, Yedong, et al.
Veröffentlicht: (2026)
von: Shen, Yedong, et al.
Veröffentlicht: (2026)
Temporal Aware Pruning for Efficient Diffusion-based Video Generation
von: Li, Sheng, et al.
Veröffentlicht: (2026)
von: Li, Sheng, et al.
Veröffentlicht: (2026)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
von: Zheng, Mingzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Mingzhe, et al.
Veröffentlicht: (2026)
Unified Text-Image Generation with Weakness-Targeted Post-Training
von: Chen, Jiahui, et al.
Veröffentlicht: (2026)
von: Chen, Jiahui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
von: Lu, Haocheng, et al.
Veröffentlicht: (2026) -
Diffusion Adversarial Post-Training for One-Step Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025) -
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025) -
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
von: Zhang, Bowen, et al.
Veröffentlicht: (2024) -
MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)