TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yuanze, Fan, Zhaoxin, Wang, Xinyu, Li, Gen, Qiu, Ye, Yang, Zhichao, Wu, Wenjun, Wu, Kejian, Sun, Yifan, Deng, Xiaotie, Dong, Jin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
by: Yang, Zhichao, et al.
Published: (2025)
by: Yang, Zhichao, et al.
Published: (2025)
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading
by: Hu, Yuanze, et al.
Published: (2026)
by: Hu, Yuanze, et al.
Published: (2026)
The Alignment Bottleneck
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment
by: Wu, Yusen, et al.
Published: (2026)
by: Wu, Yusen, et al.
Published: (2026)
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping
by: Wu, Yusen, et al.
Published: (2025)
by: Wu, Yusen, et al.
Published: (2025)
HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs
by: Wu, Yusen, et al.
Published: (2026)
by: Wu, Yusen, et al.
Published: (2026)
DeepRule: An Integrated Framework for Automated Business Rule Generation via Deep Predictive Modeling and Hybrid Search Optimization
by: Wu, Yusen, et al.
Published: (2025)
by: Wu, Yusen, et al.
Published: (2025)
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
by: Lan, Yuqin, et al.
Published: (2026)
by: Lan, Yuqin, et al.
Published: (2026)
On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial Attack
by: Li, Zhichao
Published: (2025)
by: Li, Zhichao
Published: (2025)
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
by: Wu, Yusen, et al.
Published: (2025)
by: Wu, Yusen, et al.
Published: (2025)
HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment
by: Meng, Ming, et al.
Published: (2025)
by: Meng, Ming, et al.
Published: (2025)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
Burger: Robust Graph Denoising-augmentation Fusion and Multi-semantic Modeling in Social Recommendation
by: Lan, Yuqin, et al.
Published: (2025)
by: Lan, Yuqin, et al.
Published: (2025)
TinyIO: Lightweight Reparameterized Inertial Odometry
by: Zhang, Shanshan, et al.
Published: (2025)
by: Zhang, Shanshan, et al.
Published: (2025)
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment
by: Li, Jiahuan, et al.
Published: (2024)
by: Li, Jiahuan, et al.
Published: (2024)
Lyapunov Probes for Hallucination Detection in Large Foundation Models
by: Luan, Bozhi, et al.
Published: (2026)
by: Luan, Bozhi, et al.
Published: (2026)
AlignFreeNet: Is Cross-Modal Pre-Alignment Necessary? An End-to-End Alignment-Free Lightweight Network for Visible-Infrared Object Detection
by: Zhu, Dingkun, et al.
Published: (2025)
by: Zhu, Dingkun, et al.
Published: (2025)
Will AI Trade? A Computational Inversion of the No-Trade Theorem
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
How Large Language Models Need Symbolism
by: Deng, Xiaotie, et al.
Published: (2025)
by: Deng, Xiaotie, et al.
Published: (2025)
A Semantics-Based Information Distribution Framework for Large Web-Based Course Forum System
by: Chim, Hung, et al.
Published: (2008)
by: Chim, Hung, et al.
Published: (2008)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
by: Chen, Boshui, et al.
Published: (2026)
by: Chen, Boshui, et al.
Published: (2026)
EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization
by: Fan, Zhaoxin, et al.
Published: (2026)
by: Fan, Zhaoxin, et al.
Published: (2026)
Inorganic Ligands Boosted Hybrid Infrared Photodetection via Energy Level Alignment and Interface Charge Transfer
by: Yuanze Hong, et al.
Published: (2025)
by: Yuanze Hong, et al.
Published: (2025)
Extremal values of $L^2$-Pohozaev manifolds and their applications
by: Liu, Taicheng, et al.
Published: (2024)
by: Liu, Taicheng, et al.
Published: (2024)
Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL
by: Wu, Ruitao, et al.
Published: (2025)
by: Wu, Ruitao, et al.
Published: (2025)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
by: Zhou, Xukun, et al.
Published: (2024)
by: Zhou, Xukun, et al.
Published: (2024)
FuRPE: Learning Full-body Reconstruction from Part Experts
by: Fan, Zhaoxin, et al.
Published: (2022)
by: Fan, Zhaoxin, et al.
Published: (2022)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
by: Wu, Yuhang, et al.
Published: (2024)
by: Wu, Yuhang, et al.
Published: (2024)
EMG-UP: Unsupervised Personalization in Cross-User EMG Gesture Recognition
by: Wang, Nana, et al.
Published: (2025)
by: Wang, Nana, et al.
Published: (2025)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
by: Lu, Keming, et al.
Published: (2024)
by: Lu, Keming, et al.
Published: (2024)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
Aligned Bamboo Fiber‐Induced Crystallinity Mitigation of Lightweight Polymer Composite Enables Ultrahigh Strength and Unprecedented Toughness
by: Jingkun Hou, et al.
Published: (2024)
by: Jingkun Hou, et al.
Published: (2024)
RoboPARA: Dual-Arm Robot Planning with Parallel Allocation and Recomposition Across Tasks
by: Duan, Shiying, et al.
Published: (2025)
by: Duan, Shiying, et al.
Published: (2025)
VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck
by: Zhang, Feiran, et al.
Published: (2026)
by: Zhang, Feiran, et al.
Published: (2026)
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
Similar Items
-
Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
by: Yang, Zhichao, et al.
Published: (2025) -
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading
by: Hu, Yuanze, et al.
Published: (2026) -
The Alignment Bottleneck
by: Cao, Wenjun
Published: (2025) -
MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment
by: Wu, Yusen, et al.
Published: (2026) -
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
by: Zhang, Xingjian, et al.
Published: (2025)