Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
Fuente:
arXiv
Saved in:
| Main Authors: | Cuong, Dinh Viet, Le, Hoang-Bao, Nguyen, An Pham Ngoc, Zhou, Liting, Gurrin, Cathal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025)
by: Tran, Quang-Linh, et al.
Published: (2025)
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
by: Le, Hoang-Bao, et al.
Published: (2025)
by: Le, Hoang-Bao, et al.
Published: (2025)
Results of the 2025 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
by: Tu, Jinzhe, et al.
Published: (2026)
by: Tu, Jinzhe, et al.
Published: (2026)
Results of the 2024 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2024)
by: Rossetto, Luca, et al.
Published: (2024)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
by: Nguyen, Thanh Tam, et al.
Published: (2024)
by: Nguyen, Thanh Tam, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
by: Phung, Thu Hang, et al.
Published: (2026)
by: Phung, Thu Hang, et al.
Published: (2026)
MVP: Winning Solution to SMP Challenge 2025 Video Track
by: Ye, Liliang, et al.
Published: (2025)
by: Ye, Liliang, et al.
Published: (2025)
LookupForensics: A Large-Scale Multi-Task Dataset for Multi-Phase Image-Based Fact Verification
by: Cui, Shuhan, et al.
Published: (2024)
by: Cui, Shuhan, et al.
Published: (2024)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Learning Video Context as Interleaved Multimodal Sequences
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
by: Tran, Allie, et al.
Published: (2025)
by: Tran, Allie, et al.
Published: (2025)
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
by: Chiu, Pin-Yen, et al.
Published: (2025)
by: Chiu, Pin-Yen, et al.
Published: (2025)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
by: Niu, Fuqiang, et al.
Published: (2024)
by: Niu, Fuqiang, et al.
Published: (2024)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
by: Nguyen, Cam-Van Thi, et al.
Published: (2023)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
by: Liu, Buyu, et al.
Published: (2024)
by: Liu, Buyu, et al.
Published: (2024)
TALE: Training-free Cross-domain Image Composition via Adaptive Latent Manipulation and Energy-guided Optimization
by: Pham, Kien T., et al.
Published: (2024)
by: Pham, Kien T., et al.
Published: (2024)
GTATrack: Winner Solution to SoccerTrack 2025 with Deep-EIoU and Global Tracklet Association
by: Jian, Rong-Lin, et al.
Published: (2026)
by: Jian, Rong-Lin, et al.
Published: (2026)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
by: Vo-Thanh, Hoang-Son, et al.
Published: (2024)
Born to Run, Programmed to Play: Mapping the Extended Reality Exergames Landscape
by: Karaosmanoglu, Sukran, et al.
Published: (2024)
by: Karaosmanoglu, Sukran, et al.
Published: (2024)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
by: Lu, Jinghui, et al.
Published: (2024)
by: Lu, Jinghui, et al.
Published: (2024)
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
by: Fan, Xinqi, et al.
Published: (2025)
by: Fan, Xinqi, et al.
Published: (2025)
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
by: Phan, Van-Hoang, et al.
Published: (2025)
by: Phan, Van-Hoang, et al.
Published: (2025)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
by: Zhang, Qintong, et al.
Published: (2024)
by: Zhang, Qintong, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
Scene Graph Generation with Role-Playing Large Language Models
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
by: Luu, Thanh-Danh, et al.
Published: (2025)
by: Luu, Thanh-Danh, et al.
Published: (2025)
Multimodal Sentiment Analysis Based on Causal Reasoning
by: Chen, Fuhai, et al.
Published: (2024)
by: Chen, Fuhai, et al.
Published: (2024)
Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
Similar Items
-
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
by: Tran, Quang-Linh, et al.
Published: (2025) -
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
by: Le, Hoang-Bao, et al.
Published: (2025) -
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
by: Le, Hoang-Bao, et al.
Published: (2025) -
Results of the 2025 Video Browser Showdown
by: Rossetto, Luca, et al.
Published: (2025) -
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
by: Tu, Jinzhe, et al.
Published: (2026)