Salvato in:
| Autori principali: | Dell'Erba, Samuele, Bagdanov, Andrew D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2511.20821 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
di: Niu, Yuwei, et al.
Pubblicazione: (2025)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
di: Feng, Yichen, et al.
Pubblicazione: (2026)
di: Feng, Yichen, et al.
Pubblicazione: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
di: Dua, Karan, et al.
Pubblicazione: (2025)
di: Dua, Karan, et al.
Pubblicazione: (2025)
Chat-Driven Text Generation and Interaction for Person Retrieval
di: Xie, Zequn, et al.
Pubblicazione: (2025)
di: Xie, Zequn, et al.
Pubblicazione: (2025)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
di: Skripkin, Matvey, et al.
Pubblicazione: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025)
di: Li, Danyang, et al.
Pubblicazione: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
di: Shahin, Nada, et al.
Pubblicazione: (2025)
di: Shahin, Nada, et al.
Pubblicazione: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
di: Liu, Shanyuan, et al.
Pubblicazione: (2023)
di: Liu, Shanyuan, et al.
Pubblicazione: (2023)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
di: Chen, Junting, et al.
Pubblicazione: (2024)
di: Chen, Junting, et al.
Pubblicazione: (2024)
OkanNet: A Lightweight Deep Learning Architecture for Classification of Brain Tumor from MRI Images
di: Uçar, Okan, et al.
Pubblicazione: (2026)
di: Uçar, Okan, et al.
Pubblicazione: (2026)
OmniFusion Technical Report
di: Goncharova, Elizaveta, et al.
Pubblicazione: (2024)
di: Goncharova, Elizaveta, et al.
Pubblicazione: (2024)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
di: Shahin, Nada, et al.
Pubblicazione: (2026)
di: Shahin, Nada, et al.
Pubblicazione: (2026)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
di: Xu, Zhenyu, et al.
Pubblicazione: (2025)
di: Xu, Zhenyu, et al.
Pubblicazione: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
di: Bandyopadhyay, Saptarashmi, et al.
Pubblicazione: (2025)
di: Bandyopadhyay, Saptarashmi, et al.
Pubblicazione: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
di: Seo, Huichan, et al.
Pubblicazione: (2025)
di: Seo, Huichan, et al.
Pubblicazione: (2025)
A Surveillance Based Interactive Robot
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
di: Kavimandan, Kshitij, et al.
Pubblicazione: (2025)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
di: Bonial, Claire, et al.
Pubblicazione: (2024)
di: Bonial, Claire, et al.
Pubblicazione: (2024)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
di: Lukin, Stephanie M., et al.
Pubblicazione: (2024)
di: Lukin, Stephanie M., et al.
Pubblicazione: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
di: Mahdian, Navid, et al.
Pubblicazione: (2024)
di: Mahdian, Navid, et al.
Pubblicazione: (2024)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
di: Tourani, Ali, et al.
Pubblicazione: (2023)
di: Tourani, Ali, et al.
Pubblicazione: (2023)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
di: Panek, Vojtech, et al.
Pubblicazione: (2026)
di: Panek, Vojtech, et al.
Pubblicazione: (2026)
Pro-DG: Procedural Diffusion Guidance for Architectural Facade Generation
di: Plocharski, Aleksander, et al.
Pubblicazione: (2025)
di: Plocharski, Aleksander, et al.
Pubblicazione: (2025)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
di: Annese, Luca, et al.
Pubblicazione: (2025)
di: Annese, Luca, et al.
Pubblicazione: (2025)
IntrinsiX: High-Quality PBR Generation using Image Priors
di: Kocsis, Peter, et al.
Pubblicazione: (2025)
di: Kocsis, Peter, et al.
Pubblicazione: (2025)
GroundCap: A Visually Grounded Image Captioning Dataset
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
di: Portelance, Eva, et al.
Pubblicazione: (2023)
di: Portelance, Eva, et al.
Pubblicazione: (2023)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
di: Panek, Vojtech, et al.
Pubblicazione: (2024)
di: Panek, Vojtech, et al.
Pubblicazione: (2024)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
di: Bucher, Martin JJ., et al.
Pubblicazione: (2025)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
di: Cao, Jingtao, et al.
Pubblicazione: (2024)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
di: Masrourisaadat, Nila, et al.
Pubblicazione: (2024)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
di: Deichler, Anna, et al.
Pubblicazione: (2025)
di: Deichler, Anna, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
di: Niu, Yuwei, et al.
Pubblicazione: (2025) -
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
di: Feng, Yichen, et al.
Pubblicazione: (2026) -
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
di: Dua, Karan, et al.
Pubblicazione: (2025) -
Chat-Driven Text Generation and Interaction for Person Retrieval
di: Xie, Zequn, et al.
Pubblicazione: (2025)