The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Serra, Alessandro Pietro, Ortu, Francesco, Panizon, Emanuele, Valeriani, Lucrezia, Basile, Lorenzo, Ansuini, Alessio, Doimo, Diego, Cazzaniga, Alberto |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The representation landscape of few-shot learning and fine-tuning in large language models
por: Doimo, Diego, et al.
Publicado: (2024)
por: Doimo, Diego, et al.
Publicado: (2024)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
por: Ortu, Francesco, et al.
Publicado: (2025)
por: Ortu, Francesco, et al.
Publicado: (2025)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
por: Basile, Lorenzo, et al.
Publicado: (2025)
por: Basile, Lorenzo, et al.
Publicado: (2025)
Emergent representations in networks trained with the Forward-Forward algorithm
por: Tosato, Niccolò, et al.
Publicado: (2023)
por: Tosato, Niccolò, et al.
Publicado: (2023)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
por: Ortu, Francesco, et al.
Publicado: (2024)
por: Ortu, Francesco, et al.
Publicado: (2024)
Optimal Control of a Mesoscopic Information Engine
por: Panizon, Emanuele
Publicado: (2026)
por: Panizon, Emanuele
Publicado: (2026)
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
por: Cazzaniga, Luca
Publicado: (2026)
por: Cazzaniga, Luca
Publicado: (2026)
Blending adversarial training and representation-conditional purification via aggregation improves adversarial robustness
por: Ballarin, Emanuele, et al.
Publicado: (2023)
por: Ballarin, Emanuele, et al.
Publicado: (2023)
Persistent Topological Features in Large Language Models
por: Gardinazzi, Yuri, et al.
Publicado: (2024)
por: Gardinazzi, Yuri, et al.
Publicado: (2024)
Olfactory search
por: Celani, Antonio, et al.
Publicado: (2024)
por: Celani, Antonio, et al.
Publicado: (2024)
An unsupervised tour through the hidden pathways of deep neural networks
por: Doimo, Diego
Publicado: (2025)
por: Doimo, Diego
Publicado: (2025)
Stable cooperation emerges in stochastic multiplicative growth
por: Fant, Lorenzo, et al.
Publicado: (2022)
por: Fant, Lorenzo, et al.
Publicado: (2022)
FIGURA: A Modular Prompt Engineering Method for Artistic Figure Photography in Safety-Filtered Text-to-Image Models
por: Cazzaniga, Luca
Publicado: (2026)
por: Cazzaniga, Luca
Publicado: (2026)
Interpreting and Steering Protein Language Models through Sparse Autoencoders
por: Garcia, Edith Natalia Villegas, et al.
Publicado: (2025)
por: Garcia, Edith Natalia Villegas, et al.
Publicado: (2025)
Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization
por: Alessi, Michele, et al.
Publicado: (2025)
por: Alessi, Michele, et al.
Publicado: (2025)
Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective
por: Aladrah, Nicola, et al.
Publicado: (2026)
por: Aladrah, Nicola, et al.
Publicado: (2026)
Are LLMs Good Safety Agents or a Propaganda Engine?
por: Yadav, Neemesh, et al.
Publicado: (2025)
por: Yadav, Neemesh, et al.
Publicado: (2025)
Understanding Variational Autoencoders with Intrinsic Dimension and Information Imbalance
por: Camboulin, Charles, et al.
Publicado: (2024)
por: Camboulin, Charles, et al.
Publicado: (2024)
Counterfactual rewards promote collective transport using individually controlled swarm microrobots
por: Heuthe, Veit-Lorenz, et al.
Publicado: (2024)
por: Heuthe, Veit-Lorenz, et al.
Publicado: (2024)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
por: Ortu, Francesco, et al.
Publicado: (2026)
por: Ortu, Francesco, et al.
Publicado: (2026)
Detach-ROCKET: Sequential feature selection for time series classification with random convolutional kernels
por: Uribarri, Gonzalo, et al.
Publicado: (2023)
por: Uribarri, Gonzalo, et al.
Publicado: (2023)
Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations Interpretability
por: Zhu, Zhiyu, et al.
Publicado: (2025)
por: Zhu, Zhiyu, et al.
Publicado: (2025)
ResiDual Transformer Alignment with Spectral Decomposition
por: Basile, Lorenzo, et al.
Publicado: (2024)
por: Basile, Lorenzo, et al.
Publicado: (2024)
Perceptual misalignment of texture representations in convolutional neural networks
por: de Paolis, Ludovica, et al.
Publicado: (2026)
por: de Paolis, Ludovica, et al.
Publicado: (2026)
Zigzag Persistence of Neural Responses to Time-Varying Stimuli
por: Gardinazzi, Yuri, et al.
Publicado: (2026)
por: Gardinazzi, Yuri, et al.
Publicado: (2026)
Moment maps and stability of holomorphic submersions
por: Ortu, Annamaria
Publicado: (2024)
por: Ortu, Annamaria
Publicado: (2024)
Age determination of sediment core MONDOVI, Refugio Mondovi, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Age determination of sediment core BIECAI, Torbiera del Biecai, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Pollen profile MONDOVI, Refugio Mondovi, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Age determination of sediment core FATE, Lago delle Fate, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Age determination of sediment core ORGIALS, Laghi dellOrgials, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Pollen profile FATE, Lago delle Fate, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Pollen profile BIECAI, Torbiera del Biecai, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Pollen profile ORGIALS, Laghi dellOrgials, Italy
por: Ortu, Elena
Publicado: (2010)
por: Ortu, Elena
Publicado: (2010)
Bounding Box-Guided Diffusion for Synthesizing Industrial Images and Segmentation Map
por: Caruso, Emanuele, et al.
Publicado: (2025)
por: Caruso, Emanuele, et al.
Publicado: (2025)
Evaluation of sorption isotherms in snacks with pregelatinized cassava
por: Amanda Cazzaniga
Publicado: (2022)
por: Amanda Cazzaniga
Publicado: (2022)
Usos escolares de Internet en adolescentes de sectores populares
por: Diego Basile
Publicado: (2013)
por: Diego Basile
Publicado: (2013)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
por: Tosato, Lucrezia, et al.
Publicado: (2026)
por: Tosato, Lucrezia, et al.
Publicado: (2026)
I2EDL: Interactive Instruction Error Detection and Localization
por: Taioli, Francesco, et al.
Publicado: (2024)
por: Taioli, Francesco, et al.
Publicado: (2024)
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
por: Cheng, Emily, et al.
Publicado: (2024)
por: Cheng, Emily, et al.
Publicado: (2024)
Ejemplares similares
-
The representation landscape of few-shot learning and fine-tuning in large language models
por: Doimo, Diego, et al.
Publicado: (2024) -
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
por: Ortu, Francesco, et al.
Publicado: (2025) -
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
por: Basile, Lorenzo, et al.
Publicado: (2025) -
Emergent representations in networks trained with the Forward-Forward algorithm
por: Tosato, Niccolò, et al.
Publicado: (2023) -
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
por: Ortu, Francesco, et al.
Publicado: (2024)