STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Pukun, Wang, Longxiang, Chen, Chen, Wang, Peicheng, Zhou, Fanqing, Li, Runze, Huang, Haojian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
von: Zhao, Pukun, et al.
Veröffentlicht: (2025)
von: Zhao, Pukun, et al.
Veröffentlicht: (2025)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
von: Huang, Haojian, et al.
Veröffentlicht: (2025)
Find, Fix, Reason: Context Repair for Video Reasoning
von: Huang, Haojian, et al.
Veröffentlicht: (2026)
von: Huang, Haojian, et al.
Veröffentlicht: (2026)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
von: Shang, Shuyao, et al.
Veröffentlicht: (2025)
von: Shang, Shuyao, et al.
Veröffentlicht: (2025)
InpaintDPO: Mitigating Spatial Relationship Hallucinations in Foreground-conditioned Inpainting via Diverse Preference Optimization
von: Li, Qirui, et al.
Veröffentlicht: (2025)
von: Li, Qirui, et al.
Veröffentlicht: (2025)
FFCA-Net: Stereo Image Compression via Fast Cascade Alignment of Side Information
von: Xia, Yichong, et al.
Veröffentlicht: (2023)
von: Xia, Yichong, et al.
Veröffentlicht: (2023)
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
von: Li, Zongzhao, et al.
Veröffentlicht: (2025)
von: Li, Zongzhao, et al.
Veröffentlicht: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
SAM-OCTA2: Layer Sequence OCTA Segmentation with Fine-tuned Segment Anything Model 2
von: Chen, Xinrun, et al.
Veröffentlicht: (2024)
von: Chen, Xinrun, et al.
Veröffentlicht: (2024)
EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning
von: Ma, Pengtao, et al.
Veröffentlicht: (2026)
von: Ma, Pengtao, et al.
Veröffentlicht: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
REALM: Retrospective Encoder Alignment for LFP Modeling
von: Wu, Peicheng, et al.
Veröffentlicht: (2026)
von: Wu, Peicheng, et al.
Veröffentlicht: (2026)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
BalancedDPO: Adaptive Multi-Metric Alignment
von: Tamboli, Dipesh, et al.
Veröffentlicht: (2025)
von: Tamboli, Dipesh, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Spatial Reasoning in Segmentation
von: Lin, Jiayi, et al.
Veröffentlicht: (2025)
von: Lin, Jiayi, et al.
Veröffentlicht: (2025)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
von: He, Xingqi, et al.
Veröffentlicht: (2025)
von: He, Xingqi, et al.
Veröffentlicht: (2025)
DPO: Dual-Perturbation Optimization for Test-time Adaptation in 3D Object Detection
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2024)
von: Chen, Zhuoxiao, et al.
Veröffentlicht: (2024)
GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image Prompting
von: Chen, Haodong, et al.
Veröffentlicht: (2024)
von: Chen, Haodong, et al.
Veröffentlicht: (2024)
Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
von: Li, Jinhao, et al.
Veröffentlicht: (2025)
Snake with Shifted Window: Learning to Adapt Vessel Pattern for OCTA Segmentation
von: Chen, Xinrun, et al.
Veröffentlicht: (2024)
von: Chen, Xinrun, et al.
Veröffentlicht: (2024)
Towards Robust Uncertainty-Aware Incomplete Multi-View Classification
von: Chen, Mulin, et al.
Veröffentlicht: (2024)
von: Chen, Mulin, et al.
Veröffentlicht: (2024)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
von: Chen, Jin, et al.
Veröffentlicht: (2024)
von: Chen, Jin, et al.
Veröffentlicht: (2024)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
von: Wang, Lu, et al.
Veröffentlicht: (2026)
von: Wang, Lu, et al.
Veröffentlicht: (2026)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
Gamma: Toward Generic Image Assessment with Mixture of Assessment Experts
von: Zhou, Hantao, et al.
Veröffentlicht: (2025)
von: Zhou, Hantao, et al.
Veröffentlicht: (2025)
PointRAFT: 3D deep learning for high-throughput prediction of potato tuber weight from partial point clouds
von: Blok, Pieter M., et al.
Veröffentlicht: (2025)
von: Blok, Pieter M., et al.
Veröffentlicht: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Safe and Reliable Diffusion Models via Subspace Projection
von: Chen, Huiqiang, et al.
Veröffentlicht: (2025)
von: Chen, Huiqiang, et al.
Veröffentlicht: (2025)
SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2026)
Video Object Segmentation with Dynamic Query Modulation
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
von: Zhou, Hantao, et al.
Veröffentlicht: (2024)
MedFlowSeg: Flow Matching for Medical Image Segmentation with Frequency-Aware Attention
von: Chen, Zhi, et al.
Veröffentlicht: (2026)
von: Chen, Zhi, et al.
Veröffentlicht: (2026)
AINet+: Advancing Superpixel Segmentation via Cascaded Association Implantation
von: Wang, Yaxiong, et al.
Veröffentlicht: (2021)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
von: Zhao, Pukun, et al.
Veröffentlicht: (2025) -
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025) -
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
von: Huang, Haojian, et al.
Veröffentlicht: (2025) -
Find, Fix, Reason: Context Repair for Video Reasoning
von: Huang, Haojian, et al.
Veröffentlicht: (2026) -
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
von: Chen, Haodong, et al.
Veröffentlicht: (2025)