SemCo: Toward Semantic Coherent Visual Relationship Forecasting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ou, Yangjun, Liu, Yao, Mi, Li, Chen, Zhenzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SemPT: Semantic Prompt Tuning for Vision-Language Models
von: Shi, Xiao, et al.
Veröffentlicht: (2025)
von: Shi, Xiao, et al.
Veröffentlicht: (2025)
Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images
von: AlaaEldin, Yara, et al.
Veröffentlicht: (2025)
von: AlaaEldin, Yara, et al.
Veröffentlicht: (2025)
SemGrasp: Semantic Grasp Generation via Language Aligned Discretization
von: Li, Kailin, et al.
Veröffentlicht: (2024)
von: Li, Kailin, et al.
Veröffentlicht: (2024)
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
von: Yang, Minghan, et al.
Veröffentlicht: (2026)
von: Yang, Minghan, et al.
Veröffentlicht: (2026)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
von: Zhang, Jesen, et al.
Veröffentlicht: (2025)
von: Zhang, Jesen, et al.
Veröffentlicht: (2025)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
von: Chen, Zisheng, et al.
Veröffentlicht: (2025)
von: Chen, Zisheng, et al.
Veröffentlicht: (2025)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
From a Social Cognitive Perspective: Context-aware Visual Social Relationship Recognition
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
von: Chen, Yanhui, et al.
Veröffentlicht: (2026)
von: Chen, Yanhui, et al.
Veröffentlicht: (2026)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
On The Coherence of Quantitative Evaluation of Visual Explanations
von: Vandersmissen, Benjamin, et al.
Veröffentlicht: (2023)
von: Vandersmissen, Benjamin, et al.
Veröffentlicht: (2023)
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
von: Zeng, Qinglin, et al.
Veröffentlicht: (2025)
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
von: Fan, Senran, et al.
Veröffentlicht: (2024)
von: Fan, Senran, et al.
Veröffentlicht: (2024)
Coherent Zero-Shot Visual Instruction Generation
von: Phung, Quynh, et al.
Veröffentlicht: (2024)
von: Phung, Quynh, et al.
Veröffentlicht: (2024)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection
von: Li, Wenqiao, et al.
Veröffentlicht: (2025)
von: Li, Wenqiao, et al.
Veröffentlicht: (2025)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
von: Liu, Xiwei, et al.
Veröffentlicht: (2026)
von: Liu, Xiwei, et al.
Veröffentlicht: (2026)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
von: Ding, Zijun, et al.
Veröffentlicht: (2025)
von: Ding, Zijun, et al.
Veröffentlicht: (2025)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
von: Liu, Lanmiao, et al.
Veröffentlicht: (2025)
von: Liu, Lanmiao, et al.
Veröffentlicht: (2025)
DOGR: Towards Versatile Visual Document Grounding and Referring
von: Zhou, Yinan, et al.
Veröffentlicht: (2024)
von: Zhou, Yinan, et al.
Veröffentlicht: (2024)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
von: Zheng, Yuze, et al.
Veröffentlicht: (2024)
von: Zheng, Yuze, et al.
Veröffentlicht: (2024)
Optimizing Dense Visual Predictions Through Multi-Task Coherence and Prioritization
von: Fontana, Maxime, et al.
Veröffentlicht: (2024)
von: Fontana, Maxime, et al.
Veröffentlicht: (2024)
SemSegDepth: A Combined Model for Semantic Segmentation and Depth Completion
von: Lagos, Juan Pablo, et al.
Veröffentlicht: (2022)
von: Lagos, Juan Pablo, et al.
Veröffentlicht: (2022)
Enhancing Steganographic Text Extraction: Evaluating the Impact of NLP Models on Accuracy and Semantic Coherence
von: Li, Mingyang, et al.
Veröffentlicht: (2024)
von: Li, Mingyang, et al.
Veröffentlicht: (2024)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
Towards Dynamic and Small Objects Refinement for Unsupervised Domain Adaptative Nighttime Semantic Segmentation
von: Pan, Jingyi, et al.
Veröffentlicht: (2023)
von: Pan, Jingyi, et al.
Veröffentlicht: (2023)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
Interactive Visual Assessment for Text-to-Image Generation Models
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2024)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
von: Liu, Weimin, et al.
Veröffentlicht: (2026)
von: Liu, Weimin, et al.
Veröffentlicht: (2026)
Language-Driven Visual Consensus for Zero-Shot Semantic Segmentation
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
von: Yu, Wenlong, et al.
Veröffentlicht: (2025)
von: Yu, Wenlong, et al.
Veröffentlicht: (2025)
Visual Prompt Discovery via Semantic Exploration
von: Kim, Jaechang, et al.
Veröffentlicht: (2026)
von: Kim, Jaechang, et al.
Veröffentlicht: (2026)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
von: Li, Yueying, et al.
Veröffentlicht: (2026)
von: Li, Yueying, et al.
Veröffentlicht: (2026)
SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
von: Fu, Chenhao, et al.
Veröffentlicht: (2026)
von: Fu, Chenhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SemPT: Semantic Prompt Tuning for Vision-Language Models
von: Shi, Xiao, et al.
Veröffentlicht: (2025) -
Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images
von: AlaaEldin, Yara, et al.
Veröffentlicht: (2025) -
SemGrasp: Semantic Grasp Generation via Language Aligned Discretization
von: Li, Kailin, et al.
Veröffentlicht: (2024) -
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
von: Yang, Minghan, et al.
Veröffentlicht: (2026) -
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
von: Zhang, Jesen, et al.
Veröffentlicht: (2025)