Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Vural, Hatice Merve, Kukul, Doga, Ozlu, Ege Erdem, Arikan, Demir Ekin, Mankoff, Bob, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
di: Ercan, Burak, et al.
Pubblicazione: (2024)
di: Ercan, Burak, et al.
Pubblicazione: (2024)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
di: Ekin, Yigit, et al.
Pubblicazione: (2024)
di: Ekin, Yigit, et al.
Pubblicazione: (2024)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
di: Sanli, Enes, et al.
Pubblicazione: (2025)
di: Sanli, Enes, et al.
Pubblicazione: (2025)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
di: Ercan, Burak, et al.
Pubblicazione: (2023)
di: Ercan, Burak, et al.
Pubblicazione: (2023)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
di: Karanfil, Enes, et al.
Pubblicazione: (2025)
di: Karanfil, Enes, et al.
Pubblicazione: (2025)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
di: Dogan, Mustafa, et al.
Pubblicazione: (2026)
di: Dogan, Mustafa, et al.
Pubblicazione: (2026)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
di: Dogan, Mustafa, et al.
Pubblicazione: (2024)
di: Dogan, Mustafa, et al.
Pubblicazione: (2024)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
di: Bond, Andrew, et al.
Pubblicazione: (2026)
di: Bond, Andrew, et al.
Pubblicazione: (2026)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
di: Ercan, Burak, et al.
Pubblicazione: (2023)
di: Ercan, Burak, et al.
Pubblicazione: (2023)
Sequential Compositional Generalization in Multimodal Models
di: Yagcioglu, Semih, et al.
Pubblicazione: (2024)
di: Yagcioglu, Semih, et al.
Pubblicazione: (2024)
Object and Relation Centric Representations for Push Effect Prediction
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
di: Tekden, Ahmet E., et al.
Pubblicazione: (2021)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
di: Bond, Andrew, et al.
Pubblicazione: (2025)
di: Bond, Andrew, et al.
Pubblicazione: (2025)
DeVisE: Behavioral Testing of Medical Large Language Models
di: Tagliabue, Camila Zurdo, et al.
Pubblicazione: (2025)
di: Tagliabue, Camila Zurdo, et al.
Pubblicazione: (2025)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
di: Çapuk, Hakan, et al.
Pubblicazione: (2025)
di: Çapuk, Hakan, et al.
Pubblicazione: (2025)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2025)
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2025)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
di: Cokelek, Mert, et al.
Pubblicazione: (2025)
di: Cokelek, Mert, et al.
Pubblicazione: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
di: Anees, Abdul Basit, et al.
Pubblicazione: (2024)
di: Anees, Abdul Basit, et al.
Pubblicazione: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
di: Ali, Moayed Haji, et al.
Pubblicazione: (2023)
di: Ali, Moayed Haji, et al.
Pubblicazione: (2023)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
di: Biner, Burak Can, et al.
Pubblicazione: (2024)
di: Biner, Burak Can, et al.
Pubblicazione: (2024)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
di: Kizil, Muhammed Burak, et al.
Pubblicazione: (2026)
Resolution of the cosmological constant problem by unimodular gravity and signature reversal symmetry
di: Erdem, Recai
Pubblicazione: (2026)
di: Erdem, Recai
Pubblicazione: (2026)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
A Comparison of Two Different Scales in Identifying the Risk of Falls in Leukaemia and Lymphoma Patients
di: Erdem Açikel, et al.
Pubblicazione: (2025)
di: Erdem Açikel, et al.
Pubblicazione: (2025)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
di: Li, Zhaoyang, et al.
Pubblicazione: (2025)
di: Li, Zhaoyang, et al.
Pubblicazione: (2025)
From Shock to Strategy: Quality of Life Indicators as Foundations for Post-Disaster Recovery in Türkiye
di: Erdem Ayçiçek
Pubblicazione: (2025)
di: Erdem Ayçiçek
Pubblicazione: (2025)
Fig. 8 in New data on the early stages and behaviour of the endangered species Callophrys mystaphia (Lepidoptera: Lycaenidae) and its first larval parasitoid, Cotesia sp. (Hymenoptera: Braconidae)
di: Seven, Erdem
Pubblicazione: (2022)
di: Seven, Erdem
Pubblicazione: (2022)
Gravitational particle production, the cosmological tensions and fast radio bursts
di: Erdem, Recai
Pubblicazione: (2025)
di: Erdem, Recai
Pubblicazione: (2025)
Gravitational particle production and the Hubble tension
di: Erdem, Recai
Pubblicazione: (2024)
di: Erdem, Recai
Pubblicazione: (2024)
The Politics of the Welfare State in Turkey
di: Yoruk, Erdem
Pubblicazione: (2022)
di: Yoruk, Erdem
Pubblicazione: (2022)
Curriculum development in higher education: A bibliometric analysis with insights from European and Turkish contexts
di: Erdem Aksoy
Pubblicazione: (2025)
di: Erdem Aksoy
Pubblicazione: (2025)
Migrations from the "Global South" and the Informal Economy in Turkey: Laissez passer, laissez faire?
di: Esra Erdem
Pubblicazione: (2006)
di: Esra Erdem
Pubblicazione: (2006)
Effects of intraoperative diltiazem infusion on flow changes in arterial and venous grafts in coronary artery bypass graft surgery
di: Ozan Erdem
Pubblicazione: (2015)
di: Ozan Erdem
Pubblicazione: (2015)
Age model of sediment core M77/1_416
di: Erdem, Zeynep
Pubblicazione: (2019)
di: Erdem, Zeynep
Pubblicazione: (2019)
Reliability and validity of the turkish version of the self-management scale for kidney transplant recipients
di: Çiğdem Erdem
Pubblicazione: (2022)
di: Çiğdem Erdem
Pubblicazione: (2022)
Investigating the Generalized Uncertainty Principle Effects on Hawking Radiation in Rotating Linear Dilaton Black Holes
di: Sucu, Erdem
Pubblicazione: (2024)
di: Sucu, Erdem
Pubblicazione: (2024)
Learned Dictionaries with Total Variation and Non-Negativity for Single-Cell Microscopy: Convergence Theory and Deterministic Multi-Channel Cell Feature Unification
di: Altuntac, Erdem
Pubblicazione: (2026)
di: Altuntac, Erdem
Pubblicazione: (2026)
Age model of sediment core M77/2_047-2
di: Erdem, Zeynep
Pubblicazione: (2019)
di: Erdem, Zeynep
Pubblicazione: (2019)
Age model of sediment core M77/2_050-4
di: Erdem, Zeynep
Pubblicazione: (2019)
di: Erdem, Zeynep
Pubblicazione: (2019)
Local connectedness and corporate social responsibility
di: Erdem Ucar
Pubblicazione: (2025)
di: Erdem Ucar
Pubblicazione: (2025)
Attitudes towards assistive technology among teachers working in special education and rehabilitation centres in Turkey
di: Raziye Erdem
Pubblicazione: (2024)
di: Raziye Erdem
Pubblicazione: (2024)
Documenti analoghi
-
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
di: Ercan, Burak, et al.
Pubblicazione: (2024) -
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
di: Ekin, Yigit, et al.
Pubblicazione: (2024) -
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
di: Sanli, Enes, et al.
Pubblicazione: (2025) -
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
di: Ercan, Burak, et al.
Pubblicazione: (2023) -
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
di: Karanfil, Enes, et al.
Pubblicazione: (2025)