A quantitative analysis of semantic information in deep representations of text and images
Fuente:
arXiv
Saved in:
| Main Authors: | Acevedo, Santiago, Mascaretti, Andrea, Rende, Riccardo, Mahaut, Matéo, Baroni, Marco, Laio, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised detection of semantic correlations in big data
by: Acevedo, Santiago, et al.
Published: (2024)
by: Acevedo, Santiago, et al.
Published: (2024)
Mapping of attention mechanisms to a generalized Potts model
by: Rende, Riccardo, et al.
Published: (2023)
by: Rende, Riccardo, et al.
Published: (2023)
Repetitions are not all alike: distinct mechanisms sustain repetition in language models
by: Mahaut, Matéo, et al.
Published: (2025)
by: Mahaut, Matéo, et al.
Published: (2025)
Referential communication in heterogeneous communities of pre-trained visual deep networks
by: Mahaut, Matéo, et al.
Published: (2023)
by: Mahaut, Matéo, et al.
Published: (2023)
A distributional simplicity bias in the learning dynamics of transformers
by: Rende, Riccardo, et al.
Published: (2024)
by: Rende, Riccardo, et al.
Published: (2024)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
Similarity of Processing Steps in Vision Model Representations
by: Mahaut, Matéo, et al.
Published: (2026)
by: Mahaut, Matéo, et al.
Published: (2026)
Probing the contents of semantic representations from text, behavior, and brain data using the psychNorms metabase
by: Hussain, Zak, et al.
Published: (2024)
by: Hussain, Zak, et al.
Published: (2024)
Are queries and keys always relevant? A case study on Transformer wave functions
by: Rende, Riccardo, et al.
Published: (2024)
by: Rende, Riccardo, et al.
Published: (2024)
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
by: Basile, Lorenzo, et al.
Published: (2024)
by: Basile, Lorenzo, et al.
Published: (2024)
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2024)
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2024)
Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
by: Saha, Shreya, et al.
Published: (2025)
by: Saha, Shreya, et al.
Published: (2025)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
by: Gultekin, Sinan, et al.
Published: (2023)
by: Gultekin, Sinan, et al.
Published: (2023)
Automatic feature selection and weighting in molecular systems using Differentiable Information Imbalance
by: Wild, Romina, et al.
Published: (2024)
by: Wild, Romina, et al.
Published: (2024)
Stereotypical gender actions can be extracted from Web text
by: Herdağdelen, Amaç, et al.
Published: (2025)
by: Herdağdelen, Amaç, et al.
Published: (2025)
Critical biblical studies via word frequency analysis: unveiling text authorship
by: Faigenbaum-Golovin, Shira, et al.
Published: (2024)
by: Faigenbaum-Golovin, Shira, et al.
Published: (2024)
A comparison of latent semantic analysis and correspondence analysis of document-term matrices
by: Qi, Qianqian, et al.
Published: (2021)
by: Qi, Qianqian, et al.
Published: (2021)
The representation landscape of few-shot learning and fine-tuning in large language models
by: Doimo, Diego, et al.
Published: (2024)
by: Doimo, Diego, et al.
Published: (2024)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2024)
by: Rakotonirina, Nathanaël Carraz, et al.
Published: (2024)
Towards a text-based quantitative and explainable histopathology image analysis
by: Nguyen, Anh Tien, et al.
Published: (2024)
by: Nguyen, Anh Tien, et al.
Published: (2024)
Attention-aware semantic relevance predicting Chinese sentence reading
by: Sun, Kun
Published: (2024)
by: Sun, Kun
Published: (2024)
Isolating authorship from content with semantic embeddings and contrastive learning
by: Huertas-Tato, Javier, et al.
Published: (2024)
by: Huertas-Tato, Javier, et al.
Published: (2024)
Tracing Computation Density in LLMs
by: Kervadec, Corentin, et al.
Published: (2026)
by: Kervadec, Corentin, et al.
Published: (2026)
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
by: Cheng, Emily, et al.
Published: (2024)
by: Cheng, Emily, et al.
Published: (2024)
Semantic similarity prediction is better than other semantic similarity measures
by: Herbold, Steffen
Published: (2023)
by: Herbold, Steffen
Published: (2023)
The Effect of Label Noise on the Information Content of Neural Representations
by: Umar, Ali Hussaini, et al.
Published: (2025)
by: Umar, Ali Hussaini, et al.
Published: (2025)
Topological quantification of ambiguity in semantic search
by: Barillot, Thomas Roland, et al.
Published: (2024)
by: Barillot, Thomas Roland, et al.
Published: (2024)
Machine-generated text detection prevents language model collapse
by: Drayson, George, et al.
Published: (2025)
by: Drayson, George, et al.
Published: (2025)
LLM-based feature generation from text for interpretable machine learning
by: Balek, Vojtěch, et al.
Published: (2024)
by: Balek, Vojtěch, et al.
Published: (2024)
The study of short texts in digital politics: Document aggregation for topic modeling
by: Nakka, Nitheesha, et al.
Published: (2025)
by: Nakka, Nitheesha, et al.
Published: (2025)
Extractive text summarisation of Privacy Policy documents using machine learning approaches
by: Choi, Chanwoo
Published: (2024)
by: Choi, Chanwoo
Published: (2024)
AIDetx: a compression-based method for identification of machine-learning generated text
by: Almeida, Leonardo, et al.
Published: (2024)
by: Almeida, Leonardo, et al.
Published: (2024)
Cropping outperforms dropout as an augmentation strategy for self-supervised training of text embeddings
by: González-Márquez, Rita, et al.
Published: (2025)
by: González-Márquez, Rita, et al.
Published: (2025)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
by: Gao, Peiheng, et al.
Published: (2025)
by: Gao, Peiheng, et al.
Published: (2025)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
by: Feng, Jiarui, et al.
Published: (2024)
by: Feng, Jiarui, et al.
Published: (2024)
Benchmarking pre-trained text embedding models in aligning built asset information
by: Shahinmoghadam, Mehrzad, et al.
Published: (2024)
by: Shahinmoghadam, Mehrzad, et al.
Published: (2024)
Polyatomic Complexes: A topologically-informed learning representation for atomistic systems
by: Khorana, Rahul, et al.
Published: (2024)
by: Khorana, Rahul, et al.
Published: (2024)
Discovering influential text using convolutional neural networks
by: Ayers, Megan, et al.
Published: (2024)
by: Ayers, Megan, et al.
Published: (2024)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
by: Bourne, Jonathan
Published: (2025)
by: Bourne, Jonathan
Published: (2025)
Similar Items
-
Unsupervised detection of semantic correlations in big data
by: Acevedo, Santiago, et al.
Published: (2024) -
Mapping of attention mechanisms to a generalized Potts model
by: Rende, Riccardo, et al.
Published: (2023) -
Repetitions are not all alike: distinct mechanisms sustain repetition in language models
by: Mahaut, Matéo, et al.
Published: (2025) -
Referential communication in heterogeneous communities of pre-trained visual deep networks
by: Mahaut, Matéo, et al.
Published: (2023) -
A distributional simplicity bias in the learning dynamics of transformers
by: Rende, Riccardo, et al.
Published: (2024)