Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hongbo, Lu, Junyu, Han, Yan, Ma, Kai, Yang, Liang, Lin, Hongfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
von: Wang, Hongbo, et al.
Veröffentlicht: (2024)
von: Wang, Hongbo, et al.
Veröffentlicht: (2024)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
von: Lu, Xingyuan, et al.
Veröffentlicht: (2025)
von: Lu, Xingyuan, et al.
Veröffentlicht: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
von: Zheng, Huan, et al.
Veröffentlicht: (2026)
von: Zheng, Huan, et al.
Veröffentlicht: (2026)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
von: Li, Yongqi, et al.
Veröffentlicht: (2025)
LaRe: Latent Refocusing for Multimodal Reasoning
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
von: Wang, Yan, et al.
Veröffentlicht: (2024)
von: Wang, Yan, et al.
Veröffentlicht: (2024)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
von: Whitehead, Spencer, et al.
Veröffentlicht: (2024)
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
von: Chen, Meiqi, et al.
Veröffentlicht: (2024)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
Aesthetic Assessment of Chinese Handwritings Based on Vision Language Models
von: Zheng, Chen, et al.
Veröffentlicht: (2026)
von: Zheng, Chen, et al.
Veröffentlicht: (2026)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
von: Yang, Zongxin, et al.
Veröffentlicht: (2024)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
von: Xu, Hongshen, et al.
Veröffentlicht: (2024)
Toward Multimodal Conversational AI for Age-Related Macular Degeneration
von: Gu, Ran, et al.
Veröffentlicht: (2026)
von: Gu, Ran, et al.
Veröffentlicht: (2026)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue
von: Li, Zhangpu, et al.
Veröffentlicht: (2024)
von: Li, Zhangpu, et al.
Veröffentlicht: (2024)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
von: Cao, Qinglong, et al.
Veröffentlicht: (2026)
von: Cao, Qinglong, et al.
Veröffentlicht: (2026)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
von: Khandelwal, Eshika, et al.
Veröffentlicht: (2025)
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Cong-Duy, et al.
Veröffentlicht: (2025)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
von: Yang, Dingkang, et al.
Veröffentlicht: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
von: Wang, Hongbo, et al.
Veröffentlicht: (2024) -
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025) -
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
von: Yang, Hao, et al.
Veröffentlicht: (2026) -
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024) -
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)