Gespeichert in:
| Hauptverfasser: | Zhu, Xiaomeng, Wang, Changwei, Wang, Haozhe, Liu, Xinyu, Lin, Fangzhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2509.05661 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
CogDoc: Towards Unified thinking in Documents
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
Towards Scene Graph Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Thuan, et al.
Veröffentlicht: (2024)
Towards Accurate One-Stage Object Detection with AP-Loss
von: Chen, Kean, et al.
Veröffentlicht: (2019)
von: Chen, Kean, et al.
Veröffentlicht: (2019)
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
von: Wang, Changpeng, et al.
Veröffentlicht: (2025)
von: Wang, Changpeng, et al.
Veröffentlicht: (2025)
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
von: Peddi, Rohith, et al.
Veröffentlicht: (2024)
TAP-JEPA: Frozen Future-Latent Probing and Two-Stage Score Fusion for EPIC-KITCHENS-100 Action Anticipation
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
Setting the Stage: Text-Driven Scene-Consistent Image Generation
von: Xie, Cong, et al.
Veröffentlicht: (2025)
von: Xie, Cong, et al.
Veröffentlicht: (2025)
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
von: Wu, Yuqun, et al.
Veröffentlicht: (2025)
von: Wu, Yuqun, et al.
Veröffentlicht: (2025)
Anticipating Future Object Compositions without Forgetting
von: Zahran, Youssef, et al.
Veröffentlicht: (2024)
von: Zahran, Youssef, et al.
Veröffentlicht: (2024)
Anticipating Next Active Objects for Egocentric Videos
von: Thakur, Sanket, et al.
Veröffentlicht: (2023)
von: Thakur, Sanket, et al.
Veröffentlicht: (2023)
DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method
von: Zhang, Qingwen, et al.
Veröffentlicht: (2025)
von: Zhang, Qingwen, et al.
Veröffentlicht: (2025)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision
von: Wang, Yuru, et al.
Veröffentlicht: (2024)
von: Wang, Yuru, et al.
Veröffentlicht: (2024)
SynTable: A Synthetic Data Generation Pipeline for Unseen Object Amodal Instance Segmentation of Cluttered Tabletop Scenes
von: Ng, Zhili, et al.
Veröffentlicht: (2023)
von: Ng, Zhili, et al.
Veröffentlicht: (2023)
Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
von: Wu, Qirui, et al.
Veröffentlicht: (2025)
von: Wu, Qirui, et al.
Veröffentlicht: (2025)
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
von: Lin, Chenguo, et al.
Veröffentlicht: (2024)
von: Lin, Chenguo, et al.
Veröffentlicht: (2024)
MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection
von: Gao, Shuyong, et al.
Veröffentlicht: (2025)
von: Gao, Shuyong, et al.
Veröffentlicht: (2025)
World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
von: Guan, Yanchen, et al.
Veröffentlicht: (2025)
von: Guan, Yanchen, et al.
Veröffentlicht: (2025)
Anticipating Object State Changes in Long Procedural Videos
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
PEAR: Phrase-Based Hand-Object Interaction Anticipation
von: Zhang, Zichen, et al.
Veröffentlicht: (2024)
von: Zhang, Zichen, et al.
Veröffentlicht: (2024)
S^2Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
von: Pei, Jialun, et al.
Veröffentlicht: (2024)
von: Pei, Jialun, et al.
Veröffentlicht: (2024)
Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation Enhancement
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects
von: Zhou, Lei, et al.
Veröffentlicht: (2024)
von: Zhou, Lei, et al.
Veröffentlicht: (2024)
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
von: Chen, Mu, et al.
Veröffentlicht: (2025)
von: Chen, Mu, et al.
Veröffentlicht: (2025)
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
von: Pasca, Razvan-George, et al.
Veröffentlicht: (2023)
von: Pasca, Razvan-George, et al.
Veröffentlicht: (2023)
FROST-STA: Frozen Dense Features for the Ego4D Short-Term Object Interaction Anticipation
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
von: Lin, Dixuan, et al.
Veröffentlicht: (2025)
von: Lin, Dixuan, et al.
Veröffentlicht: (2025)
Object-level Scene Deocclusion
von: Liu, Zhengzhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhengzhe, et al.
Veröffentlicht: (2024)
3D Scene Generation: A Survey
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
von: Wen, Beichen, et al.
Veröffentlicht: (2025)
SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data
von: Zhang, Qingwen, et al.
Veröffentlicht: (2026)
von: Zhang, Qingwen, et al.
Veröffentlicht: (2026)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjin, et al.
Veröffentlicht: (2024)
Credible Teacher for Semi-Supervised Object Detection in Open Scene
von: Zhuang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhuang, Jingyu, et al.
Veröffentlicht: (2024)
Scene Graph Aided Radiology Report Generation
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Object Style Diffusion for Generalized Object Detection in Urban Scene
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
AP-Loss for Accurate One-Stage Object Detection
von: Chen, Kean, et al.
Veröffentlicht: (2020)
von: Chen, Kean, et al.
Veröffentlicht: (2020)
Ähnliche Einträge
-
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025) -
CogDoc: Towards Unified thinking in Documents
von: Xu, Qixin, et al.
Veröffentlicht: (2025) -
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026) -
Towards Scene Graph Anticipation
von: Peddi, Rohith, et al.
Veröffentlicht: (2024) -
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)