CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Yolo Yunlong, Zhan, Gen, Yang, Li, Liao, Yiting, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
ExpRDiff: Short-exposure Guided Diffusion Model for Realistic Local Motion Deblurring
by: Yang, Zhongbao, et al.
Published: (2024)
by: Yang, Zhongbao, et al.
Published: (2024)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
by: Tang, Yolo Yunlong, et al.
Published: (2023)
by: Tang, Yolo Yunlong, et al.
Published: (2023)
ORSIFlow: Saliency-Guided Rectified Flow for Optical Remote Sensing Salient Object Detection
by: Chen, Haojing, et al.
Published: (2026)
by: Chen, Haojing, et al.
Published: (2026)
SalFAU-Net: Saliency Fusion Attention U-Net for Salient Object Detection
by: Mulat, Kassaw Abraham, et al.
Published: (2024)
by: Mulat, Kassaw Abraham, et al.
Published: (2024)
A Saliency Enhanced Feature Fusion based multiscale RGB-D Salient Object Detection Network
by: Huang, Rui, et al.
Published: (2024)
by: Huang, Rui, et al.
Published: (2024)
TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
by: Cho, Suhwan, et al.
Published: (2025)
by: Cho, Suhwan, et al.
Published: (2025)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
by: Tang, Yolo Yunlong, et al.
Published: (2022)
by: Tang, Yolo Yunlong, et al.
Published: (2022)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
Salient Object Detection in RGB-D Videos
by: Mou, Ao, et al.
Published: (2023)
by: Mou, Ao, et al.
Published: (2023)
Salient Objects in Clutter
by: Fan, Deng-Ping, et al.
Published: (2021)
by: Fan, Deng-Ping, et al.
Published: (2021)
AIM 2024 Challenge on Video Saliency Prediction: Methods and Results
by: Moskalenko, Andrey, et al.
Published: (2024)
by: Moskalenko, Andrey, et al.
Published: (2024)
SSNet: Saliency Prior and State Space Model-based Network for Salient Object Detection in RGB-D Images
by: Panda, Gargi, et al.
Published: (2025)
by: Panda, Gargi, et al.
Published: (2025)
Scaling Concept With Text-Guided Diffusion Models
by: Huang, Chao, et al.
Published: (2024)
by: Huang, Chao, et al.
Published: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
by: Xiong, Junwen, et al.
Published: (2024)
by: Xiong, Junwen, et al.
Published: (2024)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
by: Wen, Junwei, et al.
Published: (2026)
by: Wen, Junwei, et al.
Published: (2026)
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
by: Huo, Yu, et al.
Published: (2026)
by: Huo, Yu, et al.
Published: (2026)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
FreSca: Scaling in Frequency Space Enhances Diffusion Models
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
by: Hu, Yuhang, et al.
Published: (2025)
by: Hu, Yuhang, et al.
Published: (2025)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
by: Chen, Hanmo, et al.
Published: (2026)
by: Chen, Hanmo, et al.
Published: (2026)
Robust Salient Object Detection on Compressed Images Using Convolutional Neural Networks
by: Liao, Guibiao, et al.
Published: (2024)
by: Liao, Guibiao, et al.
Published: (2024)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
TDMM-LM: Bridging Facial Understanding and Animation via Language Models
by: Song, Luchuan, et al.
Published: (2026)
by: Song, Luchuan, et al.
Published: (2026)
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
by: Shen, Yuhao, et al.
Published: (2024)
by: Shen, Yuhao, et al.
Published: (2024)
Salient Object Detection From Arbitrary Modalities
by: Huang, Nianchang, et al.
Published: (2024)
by: Huang, Nianchang, et al.
Published: (2024)
Motion-aware Memory Network for Fast Video Salient Object Detection
by: Zhao, Xing, et al.
Published: (2022)
by: Zhao, Xing, et al.
Published: (2022)
Pluralistic Salient Object Detection
by: Feng, Xuelu, et al.
Published: (2024)
by: Feng, Xuelu, et al.
Published: (2024)
SSFam: Scribble Supervised Salient Object Detection Family
by: Liu, Zhengyi, et al.
Published: (2024)
by: Liu, Zhengyi, et al.
Published: (2024)
Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
by: Zhou, Zhangjun, et al.
Published: (2024)
by: Zhou, Zhangjun, et al.
Published: (2024)
Modality Prompts for Arbitrary Modality Salient Object Detection
by: Huang, Nianchang, et al.
Published: (2024)
by: Huang, Nianchang, et al.
Published: (2024)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
Similar Items
-
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
by: Qin, Qi, et al.
Published: (2025) -
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024) -
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024) -
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025) -
ExpRDiff: Short-exposure Guided Diffusion Model for Realistic Local Motion Deblurring
by: Yang, Zhongbao, et al.
Published: (2024)