Parrot: Multilingual Visual Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Hai-Long, Zhou, Da-Wei, Li, Yang, Lu, Shiyin, Yi, Chao, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Zhan, De-Chuan, Ye, Han-Jia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Tabular Reasoning with Privileged Structured Information
von: Jiang, Jun-Peng, et al.
Veröffentlicht: (2025)
von: Jiang, Jun-Peng, et al.
Veröffentlicht: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
von: Lu, Shiyin, et al.
Veröffentlicht: (2024)
Wings: Learning Multimodal LLMs without Text-only Forgetting
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
Expandable Subspace Ensemble for Pre-Trained Model-Based Class-Incremental Learning
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
PILOT: A Pre-Trained Model-Based Continual Learning Toolbox
von: Sun, Hai-Long, et al.
Veröffentlicht: (2023)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2023)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
von: Sun, Hui, et al.
Veröffentlicht: (2025)
von: Sun, Hui, et al.
Veröffentlicht: (2025)
Continual Learning with Pre-Trained Models: A Survey
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
von: Yi, Chao, et al.
Veröffentlicht: (2024)
von: Yi, Chao, et al.
Veröffentlicht: (2024)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental Learning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2024)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
von: Yang, Wenhao, et al.
Veröffentlicht: (2025)
von: Yang, Wenhao, et al.
Veröffentlicht: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
von: Yi, Chao, et al.
Veröffentlicht: (2024)
von: Yi, Chao, et al.
Veröffentlicht: (2024)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
Task-Agnostic Guided Feature Expansion for Class-Incremental Learning
von: Zheng, Bowen, et al.
Veröffentlicht: (2025)
von: Zheng, Bowen, et al.
Veröffentlicht: (2025)
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
Addressing Imbalanced Domain-Incremental Learning through Dual-Balance Collaborative Experts
von: Li, Lan, et al.
Veröffentlicht: (2025)
von: Li, Lan, et al.
Veröffentlicht: (2025)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models
von: Sun, Yuchong, et al.
Veröffentlicht: (2023)
von: Sun, Yuchong, et al.
Veröffentlicht: (2023)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
von: Weber, Alexander Arno, et al.
Veröffentlicht: (2024)
von: Weber, Alexander Arno, et al.
Veröffentlicht: (2024)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
LayAlign: Enhancing Multilingual Reasoning in Large Language Models via Layer-Wise Adaptive Fusion and Alignment Strategy
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
TV100: A TV Series Dataset that Pre-Trained CLIP Has Not Seen
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Da-Wei, et al.
Veröffentlicht: (2024)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
von: Wu, Minghao, et al.
Veröffentlicht: (2025)
von: Wu, Minghao, et al.
Veröffentlicht: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2026)
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2026)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
von: Fu, Minghao, et al.
Veröffentlicht: (2025)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
von: Hu, JiaKui, et al.
Veröffentlicht: (2025)
Cross-Sample Relational Fusion: Unifying Domain Generalization and Class-Incremental Learning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning
von: Li, Lan, et al.
Veröffentlicht: (2025)
von: Li, Lan, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
von: Chen, Pinzhen, et al.
Veröffentlicht: (2024)
von: Chen, Pinzhen, et al.
Veröffentlicht: (2024)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
von: Shaham, Uri, et al.
Veröffentlicht: (2024)
von: Shaham, Uri, et al.
Veröffentlicht: (2024)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
von: Yang, Wenhao, et al.
Veröffentlicht: (2026)
von: Yang, Wenhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multimodal Tabular Reasoning with Privileged Structured Information
von: Jiang, Jun-Peng, et al.
Veröffentlicht: (2025) -
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
von: Lu, Shiyin, et al.
Veröffentlicht: (2024) -
Wings: Learning Multimodal LLMs without Text-only Forgetting
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024) -
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025) -
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
von: Wang, Yibo, et al.
Veröffentlicht: (2026)