Saved in:
| Main Authors: | Kong, Lingcheng, Wei, Jiateng, Shen, Hanzhang, Wang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.07356 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
by: Bai, Haolei, et al.
Published: (2026)
by: Bai, Haolei, et al.
Published: (2026)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025)
by: Surikuchi, Aditya K, et al.
Published: (2025)
Data Redaction from Conditional Generative Models
by: Kong, Zhifeng, et al.
Published: (2023)
by: Kong, Zhifeng, et al.
Published: (2023)
Words That Make Language Models Perceive
by: Wang, Sophie L., et al.
Published: (2025)
by: Wang, Sophie L., et al.
Published: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
by: Zhao, Pu, et al.
Published: (2025)
by: Zhao, Pu, et al.
Published: (2025)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
R2Gen-Mamba: A Selective State Space Model for Radiology Report Generation
by: Sun, Yongheng, et al.
Published: (2024)
by: Sun, Yongheng, et al.
Published: (2024)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
by: Mansour, Malak, et al.
Published: (2025)
by: Mansour, Malak, et al.
Published: (2025)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
by: Shen, Huawen, et al.
Published: (2024)
by: Shen, Huawen, et al.
Published: (2024)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
State Space Model for New-Generation Network Alternative to Transformers: A Survey
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Generalization in Online Reinforcement Learning for Mobile Agents
by: Gu, Li, et al.
Published: (2026)
by: Gu, Li, et al.
Published: (2026)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
by: Han, Xiaochuang, et al.
Published: (2024)
by: Han, Xiaochuang, et al.
Published: (2024)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
by: Xiao, Junfei, et al.
Published: (2023)
by: Xiao, Junfei, et al.
Published: (2023)
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
by: Chen, Menglan, et al.
Published: (2025)
by: Chen, Menglan, et al.
Published: (2025)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
by: Wen, Xin, et al.
Published: (2024)
by: Wen, Xin, et al.
Published: (2024)
On Structured State-Space Duality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
by: Yang, Cheng-Fu, et al.
Published: (2024)
by: Yang, Cheng-Fu, et al.
Published: (2024)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
by: Tang, Siao, et al.
Published: (2025)
by: Tang, Siao, et al.
Published: (2025)
DreamReward: Text-to-3D Generation with Human Preference
by: Ye, Junliang, et al.
Published: (2024)
by: Ye, Junliang, et al.
Published: (2024)
Towards Better Multi-head Attention via Channel-wise Sample Permutation
by: Yuan, Shen, et al.
Published: (2024)
by: Yuan, Shen, et al.
Published: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
A Language Anchor-Guided Method for Robust Noisy Domain Generalization
by: Dai, Zilin, et al.
Published: (2025)
by: Dai, Zilin, et al.
Published: (2025)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
by: Jung, Hoin, et al.
Published: (2026)
by: Jung, Hoin, et al.
Published: (2026)
MatMamba: A Matryoshka State Space Model
by: Shukla, Abhinav, et al.
Published: (2024)
by: Shukla, Abhinav, et al.
Published: (2024)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
by: Li, Kaican, et al.
Published: (2025)
by: Li, Kaican, et al.
Published: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
Unleashing the Potential of Model Bias for Generalized Category Discovery
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
Continual Learning Using a Kernel-Based Method Over Foundation Models
by: Momeni, Saleh, et al.
Published: (2024)
by: Momeni, Saleh, et al.
Published: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
by: Wu, Ning, et al.
Published: (2026)
by: Wu, Ning, et al.
Published: (2026)
Similar Items
-
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
by: Bai, Haolei, et al.
Published: (2026) -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024) -
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
by: Surikuchi, Aditya K, et al.
Published: (2025) -
Data Redaction from Conditional Generative Models
by: Kong, Zhifeng, et al.
Published: (2023) -
Words That Make Language Models Perceive
by: Wang, Sophie L., et al.
Published: (2025)