ICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Yixiao, Peng, Yingzhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
von: Chao, Dian, et al.
Veröffentlicht: (2024)
von: Chao, Dian, et al.
Veröffentlicht: (2024)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
Supervised Fine-tuning in turn Improves Visual Foundation Models
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
On The Coherence of Quantitative Evaluation of Visual Explanations
von: Vandersmissen, Benjamin, et al.
Veröffentlicht: (2023)
von: Vandersmissen, Benjamin, et al.
Veröffentlicht: (2023)
Ovis-Image Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
Qwen3-VL Technical Report
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
von: Bai, Shuai, et al.
Veröffentlicht: (2025)
Seed1.5-VL Technical Report
von: Guo, Dong, et al.
Veröffentlicht: (2025)
von: Guo, Dong, et al.
Veröffentlicht: (2025)
iFlyBot-VLA Technical Report
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
von: Mishra, Sandeep, et al.
Veröffentlicht: (2026)
von: Mishra, Sandeep, et al.
Veröffentlicht: (2026)
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
GR-3 Technical Report
von: Cheang, Chilam, et al.
Veröffentlicht: (2025)
von: Cheang, Chilam, et al.
Veröffentlicht: (2025)
Ovis-U1 Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
Rethinking Visual Counterfactual Explanations Through Region Constraint
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
von: Sobieski, Bartlomiej, et al.
Veröffentlicht: (2024)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
Red Teaming Visual Language Models
von: Li, Mukai, et al.
Veröffentlicht: (2024)
von: Li, Mukai, et al.
Veröffentlicht: (2024)
The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
von: Peng, Jinghan, et al.
Veröffentlicht: (2025)
von: Peng, Jinghan, et al.
Veröffentlicht: (2025)
Technical Report: Competition Solution For Modelscope-Sora
von: Chen, Shengfu, et al.
Veröffentlicht: (2024)
von: Chen, Shengfu, et al.
Veröffentlicht: (2024)
Dolphin v1.0 Technical Report
von: Weng, Taohan, et al.
Veröffentlicht: (2025)
von: Weng, Taohan, et al.
Veröffentlicht: (2025)
Motif-Video 2B: Technical Report
von: Lim, Junghwan, et al.
Veröffentlicht: (2026)
von: Lim, Junghwan, et al.
Veröffentlicht: (2026)
V-CECE: Visual Counterfactual Explanations via Conceptual Edits
von: Spanos, Nikolaos, et al.
Veröffentlicht: (2025)
von: Spanos, Nikolaos, et al.
Veröffentlicht: (2025)
Reliable or Deceptive? Investigating Gated Features for Smooth Visual Explanations in CNNs
von: Mitra, Soham, et al.
Veröffentlicht: (2024)
von: Mitra, Soham, et al.
Veröffentlicht: (2024)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ziheng, et al.
Veröffentlicht: (2025)
XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography
von: Luo, Haozhe, et al.
Veröffentlicht: (2025)
von: Luo, Haozhe, et al.
Veröffentlicht: (2025)
Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
von: Chung, Minjae, et al.
Veröffentlicht: (2025)
von: Chung, Minjae, et al.
Veröffentlicht: (2025)
ZAYA1-VL-8B Technical Report
von: Shapourian, Hassan, et al.
Veröffentlicht: (2026)
von: Shapourian, Hassan, et al.
Veröffentlicht: (2026)
PLaMo 2.1-VL Technical Report
von: Kerola, Tommi, et al.
Veröffentlicht: (2026)
von: Kerola, Tommi, et al.
Veröffentlicht: (2026)
Baichuan-Omni Technical Report
von: Li, Yadong, et al.
Veröffentlicht: (2024)
von: Li, Yadong, et al.
Veröffentlicht: (2024)
MedGemma Technical Report
von: Sellergren, Andrew, et al.
Veröffentlicht: (2025)
von: Sellergren, Andrew, et al.
Veröffentlicht: (2025)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
von: Pang, Wei, et al.
Veröffentlicht: (2024)
von: Pang, Wei, et al.
Veröffentlicht: (2024)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
von: Leem, Saebom, et al.
Veröffentlicht: (2024)
von: Leem, Saebom, et al.
Veröffentlicht: (2024)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Phi-4-reasoning-vision-15B Technical Report
von: Aneja, Jyoti, et al.
Veröffentlicht: (2026)
von: Aneja, Jyoti, et al.
Veröffentlicht: (2026)
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
Technical Note: Defining and Quantifying AND-OR Interactions for Faithful and Concise Explanation of DNNs
von: Li, Mingjie, et al.
Veröffentlicht: (2023)
von: Li, Mingjie, et al.
Veröffentlicht: (2023)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
von: Zhang, Haoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
von: Chao, Dian, et al.
Veröffentlicht: (2024) -
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024) -
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
von: Park, Se Jin, et al.
Veröffentlicht: (2024) -
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026) -
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)