Saved in:
| Main Authors: | Bevli, Aviraj, Chaybouti, Sofian, Dahou, Yasser, Hacid, Hakim, Huynh, Ngoc Dung, Khac, Phuc H. Le, Narayan, Sanath, Para, Wamiq Reyaz, Singh, Ankit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27365 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024)
by: Narayan, Sanath, et al.
Published: (2024)
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024)
by: Malartic, Quentin, et al.
Published: (2024)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)
by: Para, Wamiq Reyaz, et al.
Published: (2024)
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
by: Huynh, Ngoc Dung, et al.
Published: (2024)
by: Huynh, Ngoc Dung, et al.
Published: (2024)
Visual question answering: from early developments to recent advances -- a survey
by: Huynh, Ngoc Dung, et al.
Published: (2025)
by: Huynh, Ngoc Dung, et al.
Published: (2025)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
by: Huynh, Ngoc Dung, et al.
Published: (2025)
by: Huynh, Ngoc Dung, et al.
Published: (2025)
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
by: Chaybouti, Sofian, et al.
Published: (2020)
by: Chaybouti, Sofian, et al.
Published: (2020)
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
by: Chaybouti, Sofian, et al.
Published: (2021)
by: Chaybouti, Sofian, et al.
Published: (2021)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
by: Khac, Phúc H. Le, et al.
Published: (2024)
by: Khac, Phúc H. Le, et al.
Published: (2024)
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
by: Yagoubi, Mouadh, et al.
Published: (2025)
by: Yagoubi, Mouadh, et al.
Published: (2025)
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
by: Zuo, Jingwei, et al.
Published: (2025)
by: Zuo, Jingwei, et al.
Published: (2025)
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
by: Tran, Kim Hoang, et al.
Published: (2024)
by: Tran, Kim Hoang, et al.
Published: (2024)
GenFlow: Interactive Modular System for Image Generation
by: Nguyen, Duc-Hung, et al.
Published: (2025)
by: Nguyen, Duc-Hung, et al.
Published: (2025)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025)
by: Pham, Phuc, et al.
Published: (2025)
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
PORT: Preference Optimization on Reasoning Traces
by: Lahlou, Salem, et al.
Published: (2024)
by: Lahlou, Salem, et al.
Published: (2024)
Transformer Runtime Lab
by: Khare, Aviraj
Published: (2026)
by: Khare, Aviraj
Published: (2026)
ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification
by: Pham-Ngoc, Hai, et al.
Published: (2025)
by: Pham-Ngoc, Hai, et al.
Published: (2025)
Perception with Guarantees: Certified Pose Estimation via Reachability Analysis
by: Ladner, Tobias, et al.
Published: (2026)
by: Ladner, Tobias, et al.
Published: (2026)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
Virtual Fusion with Contrastive Learning for Single Sensor-based Activity Recognition
by: Nguyen, Duc-Anh, et al.
Published: (2023)
by: Nguyen, Duc-Anh, et al.
Published: (2023)
Phân tích xu hướng mở rộng bề mặt không thấm trong quá trình đô thị hoá tại phường Long An bằng ảnh Sentinel 2 giai đoạn 2016-2025
by: Nhân, Nguyễn Trọng, et al.
Published: (2026)
by: Nhân, Nguyễn Trọng, et al.
Published: (2026)
Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges
by: Nguyen-Le, Hong-Hanh, et al.
Published: (2026)
by: Nguyen-Le, Hong-Hanh, et al.
Published: (2026)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
by: Quang, Ngoc Bui Lam, et al.
Published: (2025)
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
Enhancing Video Summarization with Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Interactive Interface For Semantic Segmentation Dataset Synthesis
by: Tran, Ngoc-Do, et al.
Published: (2025)
by: Tran, Ngoc-Do, et al.
Published: (2025)
Similar Items
-
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025) -
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025) -
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025) -
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024) -
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024)