ARPA: A Novel Hybrid Model for Advancing Visual Word Disambiguation Using Large Language Models and Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Papastavrou, Aristi, Lymperaiou, Maria, Stamou, Giorgos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023)
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023)
The Contribution of Knowledge in Visiolinguistic Learning: A Survey on Tasks and Challenges
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
Language Models as Knowledge Bases for Visual Word Sense Disambiguation
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023)
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
Automatic Generation of Fashion Images using Prompting in Generative Machine Learning Models
von: Argyrou, Georgia, et al.
Veröffentlicht: (2024)
von: Argyrou, Georgia, et al.
Veröffentlicht: (2024)
Fine-Grained ImageNet Classification in the Wild
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval
von: Chaidos, Nikolaos, et al.
Veröffentlicht: (2025)
von: Chaidos, Nikolaos, et al.
Veröffentlicht: (2025)
Counterfactual Edits for Generative Evaluation
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023)
Prompt2Fashion: An automatically generated fashion dataset
von: Argyrou, Georgia, et al.
Veröffentlicht: (2024)
von: Argyrou, Georgia, et al.
Veröffentlicht: (2024)
V-CECE: Visual Counterfactual Explanations via Conceptual Edits
von: Spanos, Nikolaos, et al.
Veröffentlicht: (2025)
von: Spanos, Nikolaos, et al.
Veröffentlicht: (2025)
U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations
von: Dimitriou, Angeliki, et al.
Veröffentlicht: (2026)
von: Dimitriou, Angeliki, et al.
Veröffentlicht: (2026)
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2025)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2025)
Structure Your Data: Towards Semantic Graph Counterfactuals
von: Dimitriou, Angeliki, et al.
Veröffentlicht: (2024)
von: Dimitriou, Angeliki, et al.
Veröffentlicht: (2024)
Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2026)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2026)
Designing Practical Models for Isolated Word Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation
von: Nilukshi, Shashini, et al.
Veröffentlicht: (2026)
von: Nilukshi, Shashini, et al.
Veröffentlicht: (2026)
Puzzle Solving using Reasoning of Large Language Models: A Survey
von: Giadikiaroglou, Panagiotis, et al.
Veröffentlicht: (2024)
von: Giadikiaroglou, Panagiotis, et al.
Veröffentlicht: (2024)
Conceptual Contrastive Edits in Textual and Vision-Language Retrieval
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2025)
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2025)
Evaluating Counterfactual Strategic Reasoning in Large Language Models
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
Pitfalls of Scale: Investigating the Inverse Task of Redefinition in Large Language Models
von: Stringli, Elena, et al.
Veröffentlicht: (2025)
von: Stringli, Elena, et al.
Veröffentlicht: (2025)
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
RISCORE: Enhancing In-Context Riddle Solving in Language Models through Context-Reconstructed Example Augmentation
von: Panagiotopoulos, Ioannis, et al.
Veröffentlicht: (2024)
von: Panagiotopoulos, Ioannis, et al.
Veröffentlicht: (2024)
Knowledge-Based Counterfactual Queries for Visual Question Answering
von: Stoikou, Theodoti, et al.
Veröffentlicht: (2023)
von: Stoikou, Theodoti, et al.
Veröffentlicht: (2023)
AILS-NTUA at SemEval-2024 Task 9: Cracking Brain Teasers: Transformer Models for Lateral Thinking Puzzles
von: Panagiotopoulos, Ioannis, et al.
Veröffentlicht: (2024)
von: Panagiotopoulos, Ioannis, et al.
Veröffentlicht: (2024)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
Exploring Advanced Large Language Models with LLMsuite
von: Roffo, Giorgio
Veröffentlicht: (2024)
von: Roffo, Giorgio
Veröffentlicht: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
Enhancing adversarial robustness in Natural Language Inference using explanations
von: Koulakos, Alexandros, et al.
Veröffentlicht: (2024)
von: Koulakos, Alexandros, et al.
Veröffentlicht: (2024)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
von: Uehara, Kohei, et al.
Veröffentlicht: (2024)
Task-Aware Resolution Optimization for Visual Large Language Models
von: Luo, Weiqing, et al.
Veröffentlicht: (2025)
von: Luo, Weiqing, et al.
Veröffentlicht: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
von: He, Jun-Yan, et al.
Veröffentlicht: (2024)
von: He, Jun-Yan, et al.
Veröffentlicht: (2024)
AILS-NTUA at SemEval-2025 Task 3: Leveraging Large Language Models and Translation Strategies for Multilingual Hallucination Detection
von: Karkani, Dimitra, et al.
Veröffentlicht: (2025)
von: Karkani, Dimitra, et al.
Veröffentlicht: (2025)
Large Vision-Language Models for Remote Sensing Visual Question Answering
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements
von: Raptopoulos, Petros, et al.
Veröffentlicht: (2025)
von: Raptopoulos, Petros, et al.
Veröffentlicht: (2025)
Words That Make Language Models Perceive
von: Wang, Sophie L., et al.
Veröffentlicht: (2025)
von: Wang, Sophie L., et al.
Veröffentlicht: (2025)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023) -
The Contribution of Knowledge in Visiolinguistic Learning: A Survey on Tasks and Challenges
von: Lymperaiou, Maria, et al.
Veröffentlicht: (2023) -
Language Models as Knowledge Bases for Visual Word Sense Disambiguation
von: Kritharoula, Anastasia, et al.
Veröffentlicht: (2023) -
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024) -
Automatic Generation of Fashion Images using Prompting in Generative Machine Learning Models
von: Argyrou, Georgia, et al.
Veröffentlicht: (2024)