Interfacing Foundation Models' Embeddings
Fuente:
arXiv
Salvato in:
| Autori principali: | Zou, Xueyan, Li, Linjie, Wang, Jianfeng, Yang, Jianwei, Ding, Mingyu, Wei, Junyi, Yang, Zhengyuan, Li, Feng, Zhang, Hao, Liu, Shilong, Aravinthan, Arul, Lee, Yong Jae, Wang, Lijuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bring Metric Functions into Diffusion Models
di: An, Jie, et al.
Pubblicazione: (2024)
di: An, Jie, et al.
Pubblicazione: (2024)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
di: Yan, An, et al.
Pubblicazione: (2024)
di: Yan, An, et al.
Pubblicazione: (2024)
Computer-Use Agents as Judges for Generative User Interface
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
LiVOS: Light Video Object Segmentation with Gated Linear Matching
di: Liu, Qin, et al.
Pubblicazione: (2024)
di: Liu, Qin, et al.
Pubblicazione: (2024)
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
di: Wang, Alex Jinpeng, et al.
Pubblicazione: (2025)
di: Wang, Alex Jinpeng, et al.
Pubblicazione: (2025)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
di: Yu, Weihao, et al.
Pubblicazione: (2023)
di: Yu, Weihao, et al.
Pubblicazione: (2023)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
di: Yang, Zhengyuan, et al.
Pubblicazione: (2023)
di: Yang, Zhengyuan, et al.
Pubblicazione: (2023)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
di: Jain, Jitesh, et al.
Pubblicazione: (2024)
di: Jain, Jitesh, et al.
Pubblicazione: (2024)
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
Matryoshka Multimodal Models
di: Cai, Mu, et al.
Pubblicazione: (2024)
di: Cai, Mu, et al.
Pubblicazione: (2024)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
di: Wang, Alex Jinpeng, et al.
Pubblicazione: (2024)
di: Wang, Alex Jinpeng, et al.
Pubblicazione: (2024)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
di: Yu, Weihao, et al.
Pubblicazione: (2024)
di: Yu, Weihao, et al.
Pubblicazione: (2024)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
di: Zhai, Yuanhao, et al.
Pubblicazione: (2024)
di: Zhai, Yuanhao, et al.
Pubblicazione: (2024)
GenXD: Generating Any 3D and 4D Scenes
di: Zhao, Yuyang, et al.
Pubblicazione: (2024)
di: Zhao, Yuyang, et al.
Pubblicazione: (2024)
Real Deep Research for AI, Robotics and Beyond
di: Zou, Xueyan, et al.
Pubblicazione: (2025)
di: Zou, Xueyan, et al.
Pubblicazione: (2025)
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
di: Zheng, Kaizhi, et al.
Pubblicazione: (2024)
di: Zheng, Kaizhi, et al.
Pubblicazione: (2024)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
di: Liu, Fuxiao, et al.
Pubblicazione: (2023)
SITE: towards Spatial Intelligence Thorough Evaluation
di: Wang, Wenqi, et al.
Pubblicazione: (2025)
di: Wang, Wenqi, et al.
Pubblicazione: (2025)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
di: Zhai, Yuanhao, et al.
Pubblicazione: (2024)
di: Zhai, Yuanhao, et al.
Pubblicazione: (2024)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
di: Ni, Minheng, et al.
Pubblicazione: (2025)
di: Ni, Minheng, et al.
Pubblicazione: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
di: He, Xuehai, et al.
Pubblicazione: (2024)
di: He, Xuehai, et al.
Pubblicazione: (2024)
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
di: Wang, Xiyao, et al.
Pubblicazione: (2025)
di: Wang, Xiyao, et al.
Pubblicazione: (2025)
Glance: Accelerating Diffusion Models with 1 Sample
di: Dong, Zhuobai, et al.
Pubblicazione: (2025)
di: Dong, Zhuobai, et al.
Pubblicazione: (2025)
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
di: Zheng, Xiangxi, et al.
Pubblicazione: (2025)
di: Zheng, Xiangxi, et al.
Pubblicazione: (2025)
Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
di: Lin, Yan-Bo, et al.
Pubblicazione: (2025)
di: Lin, Yan-Bo, et al.
Pubblicazione: (2025)
TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering
di: Mao, Dongxing, et al.
Pubblicazione: (2026)
di: Mao, Dongxing, et al.
Pubblicazione: (2026)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
Measurement of LLM's Philosophies of Human Nature
di: Ni, Minheng, et al.
Pubblicazione: (2025)
di: Ni, Minheng, et al.
Pubblicazione: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
di: Hong, Yining, et al.
Pubblicazione: (2024)
di: Hong, Yining, et al.
Pubblicazione: (2024)
Planning with the Views via Scene Self-Exploration
di: Wang, Kangrui, et al.
Pubblicazione: (2026)
di: Wang, Kangrui, et al.
Pubblicazione: (2026)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
di: Zhang, Jihai, et al.
Pubblicazione: (2025)
Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
di: Lin, Yiqi, et al.
Pubblicazione: (2025)
di: Lin, Yiqi, et al.
Pubblicazione: (2025)
DisCo: Disentangled Control for Realistic Human Dance Generation
di: Wang, Tan, et al.
Pubblicazione: (2023)
di: Wang, Tan, et al.
Pubblicazione: (2023)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2025)
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2025)
SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2025)
di: Chiang, Cheng-Han, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Bring Metric Functions into Diffusion Models
di: An, Jie, et al.
Pubblicazione: (2024) -
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
di: Yan, An, et al.
Pubblicazione: (2024) -
Computer-Use Agents as Judges for Generative User Interface
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025) -
LiVOS: Light Video Object Segmentation with Gated Linear Matching
di: Liu, Qin, et al.
Pubblicazione: (2024) -
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
di: Wang, Alex Jinpeng, et al.
Pubblicazione: (2025)