Task Formulation Matters When Learning Continually: A Case Study in Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Nikandrou, Mavina, Yu, Lu, Suglia, Alessandro, Konstas, Ioannis, Rieser, Verena |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures
by: Pantazopoulos, Georgios, et al.
Published: (2026)
by: Pantazopoulos, Georgios, et al.
Published: (2026)
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
by: Nikandrou, Malvina, et al.
Published: (2024)
by: Nikandrou, Malvina, et al.
Published: (2024)
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
by: Parekh, Amit, et al.
Published: (2024)
by: Parekh, Amit, et al.
Published: (2024)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
When Chain-of-Thought Fails, the Solution Hides in the Hidden States
by: Mehrafarin, Houman, et al.
Published: (2026)
by: Mehrafarin, Houman, et al.
Published: (2026)
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
by: Smyth, Ethan, et al.
Published: (2025)
by: Smyth, Ethan, et al.
Published: (2025)
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Voices in a Crowd: Searching for Clusters of Unique Perspectives
by: Vitsakis, Nikolas, et al.
Published: (2024)
by: Vitsakis, Nikolas, et al.
Published: (2024)
MoRFI: Monotonic Sparse Autoencoder Feature Identification
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
by: Dimakopoulos, Dimitris, et al.
Published: (2026)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
by: Akl, Ahmed, et al.
Published: (2024)
by: Akl, Ahmed, et al.
Published: (2024)
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
by: Tran, Dinh Phu, et al.
Published: (2026)
by: Tran, Dinh Phu, et al.
Published: (2026)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
AgriPath: A Systematic Exploration of Architectural Trade-offs for Crop Disease Classification
by: Mooraj, Hamza, et al.
Published: (2026)
by: Mooraj, Hamza, et al.
Published: (2026)
Electrocardiogram-Language Model for Few-Shot Question Answering with Meta Learning
by: Tang, Jialu, et al.
Published: (2024)
by: Tang, Jialu, et al.
Published: (2024)
OWLViz: An Open-World Benchmark for Visual Question Answering
by: Nguyen, Thuy, et al.
Published: (2025)
by: Nguyen, Thuy, et al.
Published: (2025)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
by: Weng, Weixi, et al.
Published: (2024)
by: Weng, Weixi, et al.
Published: (2024)
Retrieval-Augmented Generation for Domain-Specific Question Answering: A Case Study on Pittsburgh and CMU
by: Sun, Haojia, et al.
Published: (2024)
by: Sun, Haojia, et al.
Published: (2024)
Causal Question Answering with Reinforcement Learning
by: Blübaum, Lukas, et al.
Published: (2023)
by: Blübaum, Lukas, et al.
Published: (2023)
Theory and interpretability of Quantum Extreme Learning Machines: a Pauli-transfer matrix approach
by: Gross, Markus, et al.
Published: (2026)
by: Gross, Markus, et al.
Published: (2026)
PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
by: Cloete, Jacques, et al.
Published: (2026)
by: Cloete, Jacques, et al.
Published: (2026)
A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering
by: Zhang, Chenliang, et al.
Published: (2025)
by: Zhang, Chenliang, et al.
Published: (2025)
Retrieval Augmented Question Answering: When Should LLMs Admit Ignorance?
by: Wang, Dingmin, et al.
Published: (2025)
by: Wang, Dingmin, et al.
Published: (2025)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
Iterative In-Context Learning to Enhance LLMs Abstract Reasoning: The Case-Study of Algebraic Tasks
by: Fioravanti, Stefano, et al.
Published: (2025)
by: Fioravanti, Stefano, et al.
Published: (2025)
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
Exploring Diverse Methods in Visual Question Answering
by: Li, Panfeng, et al.
Published: (2024)
by: Li, Panfeng, et al.
Published: (2024)
UniPACT: A Multimodal Framework for Prognostic Question Answering on Raw ECG and Structured EHR
by: Tang, Jialu, et al.
Published: (2026)
by: Tang, Jialu, et al.
Published: (2026)
A Dataset for Spatiotemporal-Sensitive POI Question Answering
by: Han, Xiao, et al.
Published: (2025)
by: Han, Xiao, et al.
Published: (2025)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
by: Chintapatla, Ishant, et al.
Published: (2025)
by: Chintapatla, Ishant, et al.
Published: (2025)
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
by: Kong, Yaxuan, et al.
Published: (2025)
by: Kong, Yaxuan, et al.
Published: (2025)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Describe Anything Model for Visual Question Answering on Text-rich Images
by: Vu, Yen-Linh, et al.
Published: (2025)
by: Vu, Yen-Linh, et al.
Published: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
Kernel-based optimization of measurement operators for quantum reservoir computers
by: Gross, Markus, et al.
Published: (2026)
by: Gross, Markus, et al.
Published: (2026)
Noncommutative Model Selection for Data Clustering and Dimension Reduction Using Relative von Neumann Entropy
by: Guzmán-Tristán, Araceli, et al.
Published: (2024)
by: Guzmán-Tristán, Araceli, et al.
Published: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
Similar Items
-
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
by: Nikandrou, Malvina, et al.
Published: (2024) -
Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures
by: Pantazopoulos, Georgios, et al.
Published: (2026) -
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
by: Nikandrou, Malvina, et al.
Published: (2024) -
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
by: Parekh, Amit, et al.
Published: (2024) -
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
by: Pantazopoulos, Georgios, et al.
Published: (2024)