Salvato in:
| Autori principali: | Zhang, Zeliang, Pham, Phu, Zhao, Wentian, Wan, Kun, Li, Yu-Jhe, Zhou, Jianing, Miranda, Daniel, Kale, Ajinkya, Xu, Chenliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2410.06169 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
di: Deng, Shijian, et al.
Pubblicazione: (2024)
di: Deng, Shijian, et al.
Pubblicazione: (2024)
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
di: Li, Yu-Jhe, et al.
Pubblicazione: (2024)
di: Li, Yu-Jhe, et al.
Pubblicazione: (2024)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
di: Zhang, Zeliang, et al.
Pubblicazione: (2025)
di: Zhang, Zeliang, et al.
Pubblicazione: (2025)
Approximated Likelihood Ratio: A Forward-Only and Parallel Framework for Boosting Neural Network Training
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
See the Text: From Tokenization to Visual Reading
di: Xing, Ling, et al.
Pubblicazione: (2025)
di: Xing, Ling, et al.
Pubblicazione: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
An Attention‐Driven Graph Transformer With Nonlinear Modeling and Neuro‐Fuzzy Fusion for High‐Order Toxic Molecular Graph Learning
di: Phu Pham
Pubblicazione: (2026)
di: Phu Pham
Pubblicazione: (2026)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
di: Feng, Mingqian, et al.
Pubblicazione: (2024)
di: Feng, Mingqian, et al.
Pubblicazione: (2024)
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
di: Xu, Wujiang, et al.
Pubblicazione: (2025)
di: Xu, Wujiang, et al.
Pubblicazione: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
Recognize How Your Marketing Efforts May Need a Change
di: Alison Knopf
Pubblicazione: (2024)
di: Alison Knopf
Pubblicazione: (2024)
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
di: Yuan, Jiayi, et al.
Pubblicazione: (2025)
di: Yuan, Jiayi, et al.
Pubblicazione: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
di: Chen, Shimin, et al.
Pubblicazione: (2024)
di: Chen, Shimin, et al.
Pubblicazione: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
di: Kang, Seil, et al.
Pubblicazione: (2025)
di: Kang, Seil, et al.
Pubblicazione: (2025)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
di: Brack, Manuel, et al.
Pubblicazione: (2025)
di: Brack, Manuel, et al.
Pubblicazione: (2025)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
di: Chi, Donghwan, et al.
Pubblicazione: (2025)
Optimizing Crowd-Aware Multi-Agent Path Finding through Local Communication with Graph Neural Networks
di: Pham, Phu, et al.
Pubblicazione: (2023)
di: Pham, Phu, et al.
Pubblicazione: (2023)
Integrating Corpus Analysis and ChatGPT in Teaching English Collocations: A Hybrid Approach
di: Quy Huynh Phu Pham
Pubblicazione: (2025)
di: Quy Huynh Phu Pham
Pubblicazione: (2025)
Learning to Transform Dynamically for Better Adversarial Transferability
di: Zhu, Rongyi, et al.
Pubblicazione: (2024)
di: Zhu, Rongyi, et al.
Pubblicazione: (2024)
Forward Learning with Differential Privacy
di: Feng, Mingqian, et al.
Pubblicazione: (2025)
di: Feng, Mingqian, et al.
Pubblicazione: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Why Instruction-Based Unlearning Fails in Diffusion Models?
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
Will the Inclusion of Generated Data Amplify Bias Across Generations in Future Image Classification Models?
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
di: Zhao, Guangxiang, et al.
Pubblicazione: (2026)
di: Zhao, Guangxiang, et al.
Pubblicazione: (2026)
Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
di: Young, Shawn, et al.
Pubblicazione: (2025)
di: Young, Shawn, et al.
Pubblicazione: (2025)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
di: Li, Kevin, et al.
Pubblicazione: (2025)
di: Li, Kevin, et al.
Pubblicazione: (2025)
The Only Copyright Law We Need.
di: Toohey, Daniel
Pubblicazione: (1984)
di: Toohey, Daniel
Pubblicazione: (1984)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
di: Huang, Runhui, et al.
Pubblicazione: (2025)
di: Huang, Runhui, et al.
Pubblicazione: (2025)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
di: Liu, Jinming, et al.
Pubblicazione: (2025)
di: Liu, Jinming, et al.
Pubblicazione: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Targeted Forgetting of Image Subgroups in CLIP Models
di: Zhang, Zeliang, et al.
Pubblicazione: (2025)
di: Zhang, Zeliang, et al.
Pubblicazione: (2025)
Figure 4 in A new species of the stag-beetle genus Aegus MacLeay, 1819 (Coleoptera: Lucanidae) from Northern Vietnam
di: Yamamoto, Shûhei, et al.
Pubblicazione: (2025)
di: Yamamoto, Shûhei, et al.
Pubblicazione: (2025)
FIGURE 3 in Odontolabis pareoxa vietnamensis Yamamoto & Pham, a new subspecies of stag beetle (Coleoptera: Lucanidae) from northern Vietnam
di: Yamamoto, Shûhei, et al.
Pubblicazione: (2025)
di: Yamamoto, Shûhei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
di: Deng, Shijian, et al.
Pubblicazione: (2024) -
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
di: Li, Yu-Jhe, et al.
Pubblicazione: (2024) -
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
di: Wang, Zhenting, et al.
Pubblicazione: (2025) -
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
di: Li, Kevin Y., et al.
Pubblicazione: (2024) -
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)