Sequential Compositional Generalization in Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yagcioglu, Semih, İnce, Osman Batur, Erdem, Aykut, Erdem, Erkut, Elliott, Desmond, Yuret, Deniz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2024)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2024)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2026)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2026)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
DeVisE: Behavioral Testing of Medical Large Language Models
von: Tagliabue, Camila Zurdo, et al.
Veröffentlicht: (2025)
von: Tagliabue, Camila Zurdo, et al.
Veröffentlicht: (2025)
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
von: Vural, Hatice Merve, et al.
Veröffentlicht: (2026)
von: Vural, Hatice Merve, et al.
Veröffentlicht: (2026)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
von: Sanli, Enes, et al.
Veröffentlicht: (2025)
von: Sanli, Enes, et al.
Veröffentlicht: (2025)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
von: Ercan, Burak, et al.
Veröffentlicht: (2024)
von: Ercan, Burak, et al.
Veröffentlicht: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
von: Ercan, Burak, et al.
Veröffentlicht: (2023)
von: Ercan, Burak, et al.
Veröffentlicht: (2023)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
von: Karanfil, Enes, et al.
Veröffentlicht: (2025)
von: Karanfil, Enes, et al.
Veröffentlicht: (2025)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
von: Bond, Andrew, et al.
Veröffentlicht: (2026)
von: Bond, Andrew, et al.
Veröffentlicht: (2026)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
von: Ercan, Burak, et al.
Veröffentlicht: (2023)
von: Ercan, Burak, et al.
Veröffentlicht: (2023)
Neurocache: Efficient Vector Retrieval for Long-range Language Modeling
von: Safaya, Ali, et al.
Veröffentlicht: (2024)
von: Safaya, Ali, et al.
Veröffentlicht: (2024)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
von: Ekin, Yigit, et al.
Veröffentlicht: (2024)
von: Ekin, Yigit, et al.
Veröffentlicht: (2024)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
von: Bond, Andrew, et al.
Veröffentlicht: (2025)
von: Bond, Andrew, et al.
Veröffentlicht: (2025)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
von: Çapuk, Hakan, et al.
Veröffentlicht: (2025)
von: Çapuk, Hakan, et al.
Veröffentlicht: (2025)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
von: Biner, Burak Can, et al.
Veröffentlicht: (2024)
von: Biner, Burak Can, et al.
Veröffentlicht: (2024)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2025)
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2025)
Sample-efficient Integration of New Modalities into Large Language Models
von: İnce, Osman Batur, et al.
Veröffentlicht: (2025)
von: İnce, Osman Batur, et al.
Veröffentlicht: (2025)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
Locations of Characters in Narratives: Andersen and Persuasion Datasets
von: Ozyurt, Batuhan, et al.
Veröffentlicht: (2025)
von: Ozyurt, Batuhan, et al.
Veröffentlicht: (2025)
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
von: Dik, Selin, et al.
Veröffentlicht: (2025)
von: Dik, Selin, et al.
Veröffentlicht: (2025)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2024)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2024)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2026)
von: Kizil, Muhammed Burak, et al.
Veröffentlicht: (2026)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
von: Cokelek, Mert, et al.
Veröffentlicht: (2025)
von: Cokelek, Mert, et al.
Veröffentlicht: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
von: Er, Yakup Abrek, et al.
Veröffentlicht: (2025)
von: Er, Yakup Abrek, et al.
Veröffentlicht: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
von: Anees, Abdul Basit, et al.
Veröffentlicht: (2024)
von: Anees, Abdul Basit, et al.
Veröffentlicht: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
von: Ali, Moayed Haji, et al.
Veröffentlicht: (2023)
von: Ali, Moayed Haji, et al.
Veröffentlicht: (2023)
Object and Relation Centric Representations for Push Effect Prediction
von: Tekden, Ahmet E., et al.
Veröffentlicht: (2021)
von: Tekden, Ahmet E., et al.
Veröffentlicht: (2021)
GLOCON Database: Design Decisions and User Manual (v1.0)
von: Hürriyetoğlu, Ali, et al.
Veröffentlicht: (2024)
von: Hürriyetoğlu, Ali, et al.
Veröffentlicht: (2024)
Domain Specific Specialization in Low-Resource Settings: The Efficacy of Offline Response-Based Knowledge Distillation in Large Language Models
von: Aslan, Erdem, et al.
Veröffentlicht: (2026)
von: Aslan, Erdem, et al.
Veröffentlicht: (2026)
Tracking Universal Features Through Fine-Tuning and Model Merging
von: Horn, Niels, et al.
Veröffentlicht: (2024)
von: Horn, Niels, et al.
Veröffentlicht: (2024)
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
von: Zhang, Mike, et al.
Veröffentlicht: (2026)
von: Zhang, Mike, et al.
Veröffentlicht: (2026)
Which books do I like?
von: Rosenbusch, Hannes, et al.
Veröffentlicht: (2025)
von: Rosenbusch, Hannes, et al.
Veröffentlicht: (2025)
Core-based Hierarchies for Efficient GraphRAG
von: Hossain, Jakir, et al.
Veröffentlicht: (2026)
von: Hossain, Jakir, et al.
Veröffentlicht: (2026)
How Do Multilingual Language Models Remember Facts?
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
von: Fierro, Constanza, et al.
Veröffentlicht: (2024)
Efficient Learning Content Retrieval with Knowledge Injection
von: Sariturk, Batuhan, et al.
Veröffentlicht: (2024)
von: Sariturk, Batuhan, et al.
Veröffentlicht: (2024)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
Back to Bytes: Revisiting Tokenization Through UTF-8
von: Moryossef, Amit, et al.
Veröffentlicht: (2025)
von: Moryossef, Amit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2024) -
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2026) -
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024) -
DeVisE: Behavioral Testing of Medical Large Language Models
von: Tagliabue, Camila Zurdo, et al.
Veröffentlicht: (2025) -
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
von: Vural, Hatice Merve, et al.
Veröffentlicht: (2026)