X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Panagopoulou, Artemis, Xue, Le, Yu, Ning, Li, Junnan, Li, Dongxu, Joty, Shafiq, Xu, Ran, Savarese, Silvio, Xiong, Caiming, Niebles, Juan Carlos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
di: Xue, Le, et al.
Pubblicazione: (2023)
di: Xue, Le, et al.
Pubblicazione: (2023)
ViUniT: Visual Unit Tests for More Robust Visual Programming
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2024)
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2024)
BLIP3o-NEXT: Next Frontier of Native Image Generation
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2025)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
di: Masry, Ahmed, et al.
Pubblicazione: (2024)
Evaluating Vision-Language Models on Bistable Images
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2024)
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2024)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
di: Awadalla, Anas, et al.
Pubblicazione: (2024)
di: Awadalla, Anas, et al.
Pubblicazione: (2024)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
Direct Judgement Preference Optimization
di: Wang, Peifeng, et al.
Pubblicazione: (2024)
di: Wang, Peifeng, et al.
Pubblicazione: (2024)
SFR-RAG: Towards Contextually Faithful LLMs
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2024)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
di: Xue, Le, et al.
Pubblicazione: (2024)
di: Xue, Le, et al.
Pubblicazione: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2025)
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2025)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Shared Imagination: LLMs Hallucinate Alike
di: Zhou, Yilun, et al.
Pubblicazione: (2024)
di: Zhou, Yilun, et al.
Pubblicazione: (2024)
CoAct-1: Computer-using Multi-Agent System with Coding Actions
di: Song, Linxin, et al.
Pubblicazione: (2025)
di: Song, Linxin, et al.
Pubblicazione: (2025)
Future Optical Flow Prediction Improves Robot Control & Video Generation
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2026)
di: Ranasinghe, Kanchana, et al.
Pubblicazione: (2026)
MapTrace: Scalable Data Generation for Route Tracing on Maps
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2025)
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2025)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
Editing Arbitrary Propositions in LLMs without Subject Labels
di: Feigenbaum, Itai, et al.
Pubblicazione: (2024)
di: Feigenbaum, Itai, et al.
Pubblicazione: (2024)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
di: Niu, Tong, et al.
Pubblicazione: (2024)
di: Niu, Tong, et al.
Pubblicazione: (2024)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
di: Li, Jierui, et al.
Pubblicazione: (2024)
di: Li, Jierui, et al.
Pubblicazione: (2024)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
di: Li, Chuyuan, et al.
Pubblicazione: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
di: Wu, Haoning, et al.
Pubblicazione: (2024)
di: Wu, Haoning, et al.
Pubblicazione: (2024)
Unsupervised Summarization Re-ranking
di: Ravaut, Mathieu, et al.
Pubblicazione: (2022)
di: Ravaut, Mathieu, et al.
Pubblicazione: (2022)
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
di: Chen, Hailin, et al.
Pubblicazione: (2023)
di: Chen, Hailin, et al.
Pubblicazione: (2023)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
di: Liu, Ye, et al.
Pubblicazione: (2024)
di: Liu, Ye, et al.
Pubblicazione: (2024)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
di: Pang, Bo, et al.
Pubblicazione: (2025)
di: Pang, Bo, et al.
Pubblicazione: (2025)
Evaluating Psychological Safety of Large Language Models
di: Li, Xingxuan, et al.
Pubblicazione: (2022)
di: Li, Xingxuan, et al.
Pubblicazione: (2022)
Artificial intelligence and democracy: Towards digital authoritarianism or a democratic upgrade?
di: Panagopoulou, Fereniki
Pubblicazione: (2025)
di: Panagopoulou, Fereniki
Pubblicazione: (2025)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
di: Luo, Ziyang, et al.
Pubblicazione: (2025)
di: Luo, Ziyang, et al.
Pubblicazione: (2025)
WALT: Web Agents that Learn Tools
di: Prabhu, Viraj, et al.
Pubblicazione: (2025)
di: Prabhu, Viraj, et al.
Pubblicazione: (2025)
NAACL2025 Tutorial: Adaptation of Large Language Models
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
di: Ke, Zixuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
di: Xue, Le, et al.
Pubblicazione: (2023) -
ViUniT: Visual Unit Tests for More Robust Visual Programming
di: Panagopoulou, Artemis, et al.
Pubblicazione: (2024) -
BLIP3o-NEXT: Next Frontier of Native Image Generation
di: Chen, Jiuhai, et al.
Pubblicazione: (2025) -
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024) -
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)