Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Vallaeys, Théophane, Shukor, Mustafa, Cord, Matthieu, Verbeek, Jakob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
VUGEN: Visual Understanding priors for GENeration
by: Chen, Xiangyi, et al.
Published: (2025)
by: Chen, Xiangyi, et al.
Published: (2025)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025)
by: Parekh, Jayneel, et al.
Published: (2025)
What Makes Multimodal In-Context Learning Work?
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
by: Corradini, Barbara Toniella, et al.
Published: (2024)
by: Corradini, Barbara Toniella, et al.
Published: (2024)
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
Scaling Laws for Native Multimodal Models
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
by: Khayatan, Pegah, et al.
Published: (2026)
by: Khayatan, Pegah, et al.
Published: (2026)
Boosting Latent Diffusion with Perceptual Objectives
by: Berrada, Tariq, et al.
Published: (2024)
by: Berrada, Tariq, et al.
Published: (2024)
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)
by: Rahimi, Ahmad, et al.
Published: (2026)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
Better (pseudo-)labels for semi-supervised instance segmentation
by: Porcher, François, et al.
Published: (2024)
by: Porcher, François, et al.
Published: (2024)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
Pioneering Perceptual Video Fluency Assessment: A Novel Task with Benchmark Dataset and Baseline
by: Xie, Qizhi, et al.
Published: (2026)
by: Xie, Qizhi, et al.
Published: (2026)
RAP: 3D Rasterization Augmented End-to-End Planning
by: Feng, Lan, et al.
Published: (2025)
by: Feng, Lan, et al.
Published: (2025)
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension
by: Cardiel, Amaia, et al.
Published: (2024)
by: Cardiel, Amaia, et al.
Published: (2024)
GaussRender: Learning 3D Occupancy with Gaussian Rendering
by: Chambon, Loïck, et al.
Published: (2025)
by: Chambon, Loïck, et al.
Published: (2025)
Halton Scheduler For Masked Generative Image Transformer
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
by: Loiseau, Thibaut, et al.
Published: (2023)
by: Loiseau, Thibaut, et al.
Published: (2023)
Augmenting Perceptual Super-Resolution via Image Quality Predictors
by: Zhang, Fengjia, et al.
Published: (2025)
by: Zhang, Fengjia, et al.
Published: (2025)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Burst Super-Resolution with Diffusion Models for Improving Perceptual Quality
by: Tokoro, Kyotaro, et al.
Published: (2024)
by: Tokoro, Kyotaro, et al.
Published: (2024)
Annealed Winner-Takes-All for Motion Forecasting
by: Xu, Yihong, et al.
Published: (2024)
by: Xu, Yihong, et al.
Published: (2024)
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
by: Zablocki, Éloi, et al.
Published: (2024)
by: Zablocki, Éloi, et al.
Published: (2024)
PointBeV: A Sparse Approach to BeV Predictions
by: Chambon, Loick, et al.
Published: (2023)
by: Chambon, Loick, et al.
Published: (2023)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025)
by: Chambon, Loick, et al.
Published: (2025)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Action100M: A Large-scale Video Action Dataset
by: Chen, Delong, et al.
Published: (2026)
by: Chen, Delong, et al.
Published: (2026)
Enhancing Image Classification with Augmentation: Data Augmentation Techniques for Improved Image Classification
by: Kumar, Saorj, et al.
Published: (2025)
by: Kumar, Saorj, et al.
Published: (2025)
Towards image compression with perfect realism at ultra-low bitrates
by: Careil, Marlène, et al.
Published: (2023)
by: Careil, Marlène, et al.
Published: (2023)
Perceptual-GS: Scene-adaptive Perceptual Densification for Gaussian Splatting
by: Zhou, Hongbi, et al.
Published: (2025)
by: Zhou, Hongbi, et al.
Published: (2025)
Perceptual Classifiers: Detecting Generative Images using Perceptual Features
by: Durbha, Krishna Srikar, et al.
Published: (2025)
by: Durbha, Krishna Srikar, et al.
Published: (2025)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
by: Berrada, Tariq, et al.
Published: (2023)
by: Berrada, Tariq, et al.
Published: (2023)
Improved Baselines with Synchronized Encoding for Universal Medical Image Segmentation
by: Yang, Sihan, et al.
Published: (2024)
by: Yang, Sihan, et al.
Published: (2024)
Similar Items
-
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025) -
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024) -
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024) -
VUGEN: Visual Understanding priors for GENeration
by: Chen, Xiangyi, et al.
Published: (2025) -
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)