Salvato in:
| Autori principali: | Ru, Jinghan, Xie, Yuxin, Zhuang, Xianwei, Yin, Yuguo, Guo, Zhihui, Liu, Zhiming, Ren, Qianli, Zou, Yuexian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.06604 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
di: Yin, Yuguo, et al.
Pubblicazione: (2025)
di: Yin, Yuguo, et al.
Pubblicazione: (2025)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
di: Ru, Jinghan, et al.
Pubblicazione: (2026)
di: Ru, Jinghan, et al.
Pubblicazione: (2026)
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)
SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
di: Luo, Jiehui, et al.
Pubblicazione: (2025)
di: Luo, Jiehui, et al.
Pubblicazione: (2025)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
di: Fu, Siyuan, et al.
Pubblicazione: (2025)
di: Fu, Siyuan, et al.
Pubblicazione: (2025)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
di: Tang, Lexiang, et al.
Pubblicazione: (2025)
di: Tang, Lexiang, et al.
Pubblicazione: (2025)
Do we really need the Rademacher complexities?
di: Bartl, Daniel, et al.
Pubblicazione: (2025)
di: Bartl, Daniel, et al.
Pubblicazione: (2025)
On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?
di: Zanella, Maxime, et al.
Pubblicazione: (2024)
di: Zanella, Maxime, et al.
Pubblicazione: (2024)
Sequestration by the biological carbon pump: Do we really know what we are talking about?
di: Andre W. Visser
Pubblicazione: (2025)
di: Andre W. Visser
Pubblicazione: (2025)
Do we really ponder about necessity of intravenous hydration in acute bronchiolitis?
di: Sule Yıldırım
Pubblicazione: (2016)
di: Sule Yıldırım
Pubblicazione: (2016)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
di: Dkhissi, Youness, et al.
Pubblicazione: (2026)
di: Dkhissi, Youness, et al.
Pubblicazione: (2026)
Religion and spirituality in counselor education: Do we really need to talk about this?
di: Jesse Fox
Pubblicazione: (2024)
di: Jesse Fox
Pubblicazione: (2024)
HeartMuLa: A Family of Open Sourced Music Foundation Models
di: Yang, Dongchao, et al.
Pubblicazione: (2026)
di: Yang, Dongchao, et al.
Pubblicazione: (2026)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
di: Xin, Yifei, et al.
Pubblicazione: (2023)
di: Xin, Yifei, et al.
Pubblicazione: (2023)
The LHC has ruled out Supersymmetry -- really?
di: Constantin, L., et al.
Pubblicazione: (2025)
di: Constantin, L., et al.
Pubblicazione: (2025)
Data filtering methods for training language models
di: Shevchenko, Egor, et al.
Pubblicazione: (2026)
di: Shevchenko, Egor, et al.
Pubblicazione: (2026)
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
di: Yin, Yongkang, et al.
Pubblicazione: (2023)
di: Yin, Yongkang, et al.
Pubblicazione: (2023)
Friendship-paradox paradox: Do most people's friends really have more friends than they do?
di: Lee, Sang Hoon
Pubblicazione: (2025)
di: Lee, Sang Hoon
Pubblicazione: (2025)
Do Multilingual LLMs have specialized language heads?
di: Naufil, Muhammad
Pubblicazione: (2026)
di: Naufil, Muhammad
Pubblicazione: (2026)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
di: Chen, Zhanpeng, et al.
Pubblicazione: (2025)
di: Chen, Zhanpeng, et al.
Pubblicazione: (2025)
How much freedom does an effectiveness metric really have?
di: Alistair Moffat, et al.
Pubblicazione: (2024)
di: Alistair Moffat, et al.
Pubblicazione: (2024)
Do prompt positions really matter?
di: Mao, Junyu, et al.
Pubblicazione: (2023)
di: Mao, Junyu, et al.
Pubblicazione: (2023)
Conceptualizing transgender experiences in psychology: Do we have a ‘true’ gender?
di: Emma F. Jackson, et al.
Pubblicazione: (2024)
di: Emma F. Jackson, et al.
Pubblicazione: (2024)
Getting the most out of your tokenizer for pre-training and domain adaptation
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
di: Dagan, Gautier, et al.
Pubblicazione: (2024)
Universal pre-training by iterated random computation
di: Bloem, Peter
Pubblicazione: (2025)
di: Bloem, Peter
Pubblicazione: (2025)
COPD: Are we using all the tools we have?
di: A. Araújo
Pubblicazione: (2016)
di: A. Araújo
Pubblicazione: (2016)
STAR: Speech-to-Audio Generation via Representation Learning
di: Xie, Zeyu, et al.
Pubblicazione: (2025)
di: Xie, Zeyu, et al.
Pubblicazione: (2025)
Uncertainty-aware sign language video retrieval with probability distribution modeling
di: Wu, Xuan, et al.
Pubblicazione: (2024)
di: Wu, Xuan, et al.
Pubblicazione: (2024)
The mass-to-flux ratio in molecular clouds. What are we really measuring?
di: Tritsis, Aris
Pubblicazione: (2025)
di: Tritsis, Aris
Pubblicazione: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
di: Xie, Zeyu, et al.
Pubblicazione: (2025)
di: Xie, Zeyu, et al.
Pubblicazione: (2025)
Do we have to fear tax competition among "new" and "old" European countries?
di: Simon Schnyder
Pubblicazione: (2006)
di: Simon Schnyder
Pubblicazione: (2006)
Do massive neutrino states really exist?
di: Shelkovkin, Danil D., et al.
Pubblicazione: (2025)
di: Shelkovkin, Danil D., et al.
Pubblicazione: (2025)
Linking agricultural conservation to water quality outcomes in the United States at multiple scales: Do we have the information we need?
di: Laura Naslund, et al.
Pubblicazione: (2025)
di: Laura Naslund, et al.
Pubblicazione: (2025)
Tensor train methods for high-dimensional nonlinear filtering problems with correlated noise
di: Meng, Yuhua, et al.
Pubblicazione: (2026)
di: Meng, Yuhua, et al.
Pubblicazione: (2026)
What we have accomplished and what we can achieve
di: A. Morais
Pubblicazione: (2014)
di: A. Morais
Pubblicazione: (2014)
Can pre-trained language models generate titles for research papers?
di: Rehman, Tohida, et al.
Pubblicazione: (2024)
di: Rehman, Tohida, et al.
Pubblicazione: (2024)
Do we have a quantum computer? Expert perspectives on current status and future prospects
di: Doyle, Liam, et al.
Pubblicazione: (2026)
di: Doyle, Liam, et al.
Pubblicazione: (2026)
Do we have to choose between economic or environmental performance? The case of the ceramic industry cluster
di: Teresa Vallet‐Bellmunt, et al.
Pubblicazione: (2024)
di: Teresa Vallet‐Bellmunt, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025) -
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025) -
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
di: Yin, Yuguo, et al.
Pubblicazione: (2025) -
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
di: Ru, Jinghan, et al.
Pubblicazione: (2026) -
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
di: Zhuang, Xianwei, et al.
Pubblicazione: (2025)