NeoBabel: A Multilingual Open Tower for Visual Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Derakhshani, Mohammad Mahdi, Varghese, Dheeraj, Fadaee, Marzieh, Snoek, Cees G. M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Continual Hyperbolic Learning of Instances and Classes
di: Ayoughi, Melika, et al.
Pubblicazione: (2025)
di: Ayoughi, Melika, et al.
Pubblicazione: (2025)
Any-Shift Prompting for Generalization over Distributions
di: Xiao, Zehao, et al.
Pubblicazione: (2024)
di: Xiao, Zehao, et al.
Pubblicazione: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)
di: Liu, Huabin, et al.
Pubblicazione: (2025)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
di: Matişan, Răzvan-Andrei, et al.
Pubblicazione: (2025)
di: Matişan, Răzvan-Andrei, et al.
Pubblicazione: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
Vero: An Open RL Recipe for General Visual Reasoning
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
di: Bhowmik, Aritra, et al.
Pubblicazione: (2024)
di: Bhowmik, Aritra, et al.
Pubblicazione: (2024)
TULIP: Token-length Upgraded CLIP
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2024)
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2024)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
di: Rastegar, Sarah, et al.
Pubblicazione: (2024)
di: Rastegar, Sarah, et al.
Pubblicazione: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
di: Du, Yingjun, et al.
Pubblicazione: (2024)
di: Du, Yingjun, et al.
Pubblicazione: (2024)
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
di: Winata, Genta Indra, et al.
Pubblicazione: (2024)
di: Winata, Genta Indra, et al.
Pubblicazione: (2024)
Parrot: Multilingual Visual Instruction Tuning
di: Sun, Hai-Long, et al.
Pubblicazione: (2024)
di: Sun, Hai-Long, et al.
Pubblicazione: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
di: Yang, Bang, et al.
Pubblicazione: (2023)
di: Yang, Bang, et al.
Pubblicazione: (2023)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
di: Hashemi, Mohammad Abuzar, et al.
Pubblicazione: (2021)
di: Hashemi, Mohammad Abuzar, et al.
Pubblicazione: (2021)
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
di: Romero, David, et al.
Pubblicazione: (2024)
di: Romero, David, et al.
Pubblicazione: (2024)
Image-Text Relation Prediction for Multilingual Tweets
di: Rikters, Matīss, et al.
Pubblicazione: (2025)
di: Rikters, Matīss, et al.
Pubblicazione: (2025)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
di: Hu, Wenbo, et al.
Pubblicazione: (2026)
di: Hu, Wenbo, et al.
Pubblicazione: (2026)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
di: Khan, Shaharukh, et al.
Pubblicazione: (2026)
di: Khan, Shaharukh, et al.
Pubblicazione: (2026)
Translation-Enhanced Multilingual Text-to-Image Generation
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2025)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2025)
MPN: Leveraging Multilingual Patch Neuron for Cross-lingual Model Editing
di: Si, Nianwen, et al.
Pubblicazione: (2024)
di: Si, Nianwen, et al.
Pubblicazione: (2024)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
di: Geigle, Gregor, et al.
Pubblicazione: (2023)
di: Geigle, Gregor, et al.
Pubblicazione: (2023)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)
di: Chern, Ethan, et al.
Pubblicazione: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
di: Chen, Jiali, et al.
Pubblicazione: (2024)
di: Chen, Jiali, et al.
Pubblicazione: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
di: Lin, Haokun, et al.
Pubblicazione: (2025)
di: Lin, Haokun, et al.
Pubblicazione: (2025)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
di: Jang, Jihyoung, et al.
Pubblicazione: (2026)
di: Jang, Jihyoung, et al.
Pubblicazione: (2026)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
di: Willemsen, Bram, et al.
Pubblicazione: (2024)
di: Willemsen, Bram, et al.
Pubblicazione: (2024)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
di: Anugraha, David, et al.
Pubblicazione: (2025)
di: Anugraha, David, et al.
Pubblicazione: (2025)
TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models
di: Veenboer, Tim, et al.
Pubblicazione: (2025)
di: Veenboer, Tim, et al.
Pubblicazione: (2025)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
di: Lee, Jeongwoo, et al.
Pubblicazione: (2024)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
di: Wang, Junling, et al.
Pubblicazione: (2026)
di: Wang, Junling, et al.
Pubblicazione: (2026)
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
di: Chen, Haonan, et al.
Pubblicazione: (2025)
di: Chen, Haonan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Continual Hyperbolic Learning of Instances and Classes
di: Ayoughi, Melika, et al.
Pubblicazione: (2025) -
Any-Shift Prompting for Generalization over Distributions
di: Xiao, Zehao, et al.
Pubblicazione: (2024) -
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025) -
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
di: Matişan, Răzvan-Andrei, et al.
Pubblicazione: (2025) -
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)