Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
Fuente:
arXiv
Saved in:
| Main Authors: | Cioni, Dario, Berlincioni, Lorenzo, Becattini, Federico, del Bimbo, Alberto |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neuromorphic Valence and Arousal Estimation
by: Berlincioni, Lorenzo, et al.
Published: (2024)
by: Berlincioni, Lorenzo, et al.
Published: (2024)
Neuromorphic Face Analysis: a Survey
by: Becattini, Federico, et al.
Published: (2024)
by: Becattini, Federico, et al.
Published: (2024)
Spatio-temporal Transformers for Action Unit Classification with Event Cameras
by: Cultrera, Luca, et al.
Published: (2024)
by: Cultrera, Luca, et al.
Published: (2024)
Spike-TBR: a Noise Resilient Neuromorphic Event Representation
by: Magrini, Gabriele, et al.
Published: (2025)
by: Magrini, Gabriele, et al.
Published: (2025)
Neuromorphic Facial Analysis with Cross-Modal Supervision
by: Becattini, Federico, et al.
Published: (2024)
by: Becattini, Federico, et al.
Published: (2024)
FRED: The Florence RGB-Event Drone Dataset
by: Magrini, Gabriele, et al.
Published: (2025)
by: Magrini, Gabriele, et al.
Published: (2025)
Garment Attribute Manipulation with Multi-level Attention
by: Casula, Vittorio, et al.
Published: (2024)
by: Casula, Vittorio, et al.
Published: (2024)
SMEMO: Social Memory for Trajectory Forecasting
by: Marchetti, Francesco, et al.
Published: (2022)
by: Marchetti, Francesco, et al.
Published: (2022)
Drone Detection with Event Cameras
by: Magrini, Gabriele, et al.
Published: (2025)
by: Magrini, Gabriele, et al.
Published: (2025)
Deepfake detection by exploiting surface anomalies: the SurFake approach
by: Ciamarra, Andrea, et al.
Published: (2023)
by: Ciamarra, Andrea, et al.
Published: (2023)
Neuromorphic Drone Detection: an Event-RGB Multimodal Approach
by: Magrini, Gabriele, et al.
Published: (2024)
by: Magrini, Gabriele, et al.
Published: (2024)
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
by: Trippodo, Matteo, et al.
Published: (2025)
by: Trippodo, Matteo, et al.
Published: (2025)
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
by: Agnolucci, Lorenzo, et al.
Published: (2024)
by: Agnolucci, Lorenzo, et al.
Published: (2024)
Interactive Garment Recommendation with User in the Loop
by: Becattini, Federico, et al.
Published: (2024)
by: Becattini, Federico, et al.
Published: (2024)
Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
by: Cioni, Dario, et al.
Published: (2024)
by: Cioni, Dario, et al.
Published: (2024)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
Towards Retrieval-Augmented Architectures for Image Captioning
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Backward-Compatible Aligned Representations via an Orthogonal Transformation Layer
by: Ricci, Simone, et al.
Published: (2024)
by: Ricci, Simone, et al.
Published: (2024)
Stationary Representations: Optimally Approximating Compatibility and Implications for Improved Model Replacements
by: Biondi, Niccolò, et al.
Published: (2024)
by: Biondi, Niccolò, et al.
Published: (2024)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
by: Zhang, Jian, et al.
Published: (2025)
by: Zhang, Jian, et al.
Published: (2025)
Hands-Free Heritage: Automated 3D Scanning for Cultural Heritage Digitization
by: Ahmad, Javed, et al.
Published: (2025)
by: Ahmad, Javed, et al.
Published: (2025)
Prompt and Prejudice
by: Berlincioni, Lorenzo, et al.
Published: (2024)
by: Berlincioni, Lorenzo, et al.
Published: (2024)
Mitigating Negative Flips via Margin Preserving Training
by: Ricci, Simone, et al.
Published: (2025)
by: Ricci, Simone, et al.
Published: (2025)
Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
by: Yuan, Junyi, et al.
Published: (2025)
by: Yuan, Junyi, et al.
Published: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
by: Pujol-Perich, David, et al.
Published: (2026)
by: Pujol-Perich, David, et al.
Published: (2026)
PEPR: Privileged Event-based Predictive Regularization for Domain Generalization
by: Magrini, Gabriele, et al.
Published: (2026)
by: Magrini, Gabriele, et al.
Published: (2026)
EV-Flying: an Event-based Dataset for In-The-Wild Recognition of Flying Objects
by: Magrini, Gabriele, et al.
Published: (2025)
by: Magrini, Gabriele, et al.
Published: (2025)
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
by: Fonseca, Rui, et al.
Published: (2025)
by: Fonseca, Rui, et al.
Published: (2025)
Augmented Reality in Cultural Heritage: A Dual-Model Pipeline for 3D Artwork Reconstruction
by: Pannone, Daniele, et al.
Published: (2025)
by: Pannone, Daniele, et al.
Published: (2025)
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
by: Li, Jiaxuan, et al.
Published: (2023)
by: Li, Jiaxuan, et al.
Published: (2023)
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
by: Galliena, Tommaso, et al.
Published: (2026)
by: Galliena, Tommaso, et al.
Published: (2026)
RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
by: Long, Xiaosheng, et al.
Published: (2025)
by: Long, Xiaosheng, et al.
Published: (2025)
Gaussian Heritage: 3D Digitization of Cultural Heritage with Integrated Object Segmentation
by: Dahaghin, Mahtab, et al.
Published: (2024)
by: Dahaghin, Mahtab, et al.
Published: (2024)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
by: Dhawan, Aashish, et al.
Published: (2026)
by: Dhawan, Aashish, et al.
Published: (2026)
Multiview Progress Prediction of Robot Activities
by: Zoppellari, Elena, et al.
Published: (2026)
by: Zoppellari, Elena, et al.
Published: (2026)
Cultural Heritage 3D Reconstruction with Diffusion Networks
by: Jaramillo, Pablo, et al.
Published: (2024)
by: Jaramillo, Pablo, et al.
Published: (2024)
Depth-based Privileged Information for Boosting 3D Human Pose Estimation on RGB
by: Simoni, Alessandro, et al.
Published: (2024)
by: Simoni, Alessandro, et al.
Published: (2024)
Similar Items
-
Neuromorphic Valence and Arousal Estimation
by: Berlincioni, Lorenzo, et al.
Published: (2024) -
Neuromorphic Face Analysis: a Survey
by: Becattini, Federico, et al.
Published: (2024) -
Spatio-temporal Transformers for Action Unit Classification with Event Cameras
by: Cultrera, Luca, et al.
Published: (2024) -
Spike-TBR: a Noise Resilient Neuromorphic Event Representation
by: Magrini, Gabriele, et al.
Published: (2025) -
Neuromorphic Facial Analysis with Cross-Modal Supervision
by: Becattini, Federico, et al.
Published: (2024)