DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Run, Li, Yunshui, Chen, Longze, He, Wanwei, Lin, Ting-En, Liu, Ziqiang, Zhang, Lei, Song, Zikai, Xia, Xiaobo, Liu, Tongliang, Yang, Min, Hui, Binyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
von: Chen, Longze, et al.
Veröffentlicht: (2024)
von: Chen, Longze, et al.
Veröffentlicht: (2024)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
IP-MOT: Instance Prompt Learning for Cross-Domain Multi-Object Tracking
von: Luo, Run, et al.
Veröffentlicht: (2024)
von: Luo, Run, et al.
Veröffentlicht: (2024)
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
von: Li, Jiaming, et al.
Veröffentlicht: (2024)
Marathon: A Race Through the Realm of Long Context with Large Language Models
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
von: Zhang, Lei, et al.
Veröffentlicht: (2023)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
One-Shot Learning as Instruction Data Prospector for Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2023)
von: Li, Yunshui, et al.
Veröffentlicht: (2023)
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
von: Luo, Run, et al.
Veröffentlicht: (2024)
von: Luo, Run, et al.
Veröffentlicht: (2024)
DEEM: Dynamic Experienced Expert Modeling for Stance Detection
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
DiffusionTrack: Diffusion Model For Multi-Object Tracking
von: Luo, Run, et al.
Veröffentlicht: (2023)
von: Luo, Run, et al.
Veröffentlicht: (2023)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Omnimodal Dataset Distillation via High-order Proxy Alignment
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
von: Zhou, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhou, Yiwei, et al.
Veröffentlicht: (2024)
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
von: Chen, Longze, et al.
Veröffentlicht: (2025)
von: Chen, Longze, et al.
Veröffentlicht: (2025)
IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models
von: Zhang, Shaokun, et al.
Veröffentlicht: (2023)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2023)
LaVin-DiT: Large Vision Diffusion Transformer
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Wang, Zhaoqing, et al.
Veröffentlicht: (2024)
Consistent Image Layout Editing with Diffusion Models
von: Xia, Tao, et al.
Veröffentlicht: (2025)
von: Xia, Tao, et al.
Veröffentlicht: (2025)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Transferring Annotator- and Instance-dependent Transition Matrix for Learning from Crowds
von: Li, Shikun, et al.
Veröffentlicht: (2023)
von: Li, Shikun, et al.
Veröffentlicht: (2023)
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
von: Wan, Zhenchen, et al.
Veröffentlicht: (2024)
von: Wan, Zhenchen, et al.
Veröffentlicht: (2024)
Learning Ordinal Probabilistic Reward from Preferences
von: Chen, Longze, et al.
Veröffentlicht: (2026)
von: Chen, Longze, et al.
Veröffentlicht: (2026)
An Unsupervised Dialogue Topic Segmentation Model Based on Utterance Rewriting
von: Hou, Xia, et al.
Veröffentlicht: (2024)
von: Hou, Xia, et al.
Veröffentlicht: (2024)
Accelerating Parallel Diffusion Model Serving with Residual Compression
von: Luo, Jiajun, et al.
Veröffentlicht: (2025)
von: Luo, Jiajun, et al.
Veröffentlicht: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
Enhancing User-Centric Privacy Protection: An Interactive Framework through Diffusion Models and Machine Unlearning
von: Huang, Huaxi, et al.
Veröffentlicht: (2024)
von: Huang, Huaxi, et al.
Veröffentlicht: (2024)
Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching
von: Lin, Yexiong, et al.
Veröffentlicht: (2025)
von: Lin, Yexiong, et al.
Veröffentlicht: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Optimized View and Geometry Distillation from Multi-view Diffuser
von: Zhang, Youjia, et al.
Veröffentlicht: (2023)
von: Zhang, Youjia, et al.
Veröffentlicht: (2023)
Parallel Scaling Law for Language Models
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2025)
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
von: Xia, Yuchen, et al.
Veröffentlicht: (2025)
IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property
von: Wang, Qiyao, et al.
Veröffentlicht: (2025)
von: Wang, Qiyao, et al.
Veröffentlicht: (2025)
Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models
von: Xie, Zikai
Veröffentlicht: (2024)
von: Xie, Zikai
Veröffentlicht: (2024)
Enhancing Diffusion Models for Inverse Problems with Covariance-Aware Posterior Sampling
von: Hamidi, Shayan Mohajer, et al.
Veröffentlicht: (2024)
von: Hamidi, Shayan Mohajer, et al.
Veröffentlicht: (2024)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
von: Yin, Zhichao, et al.
Veröffentlicht: (2024)
von: Yin, Zhichao, et al.
Veröffentlicht: (2024)
Iterative Forward Tuning Boosts In-Context Learning in Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
von: Chen, Longze, et al.
Veröffentlicht: (2024) -
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025) -
IP-MOT: Instance Prompt Learning for Cross-Domain Multi-Object Tracking
von: Luo, Run, et al.
Veröffentlicht: (2024) -
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models
von: Li, Jiaming, et al.
Veröffentlicht: (2024) -
Marathon: A Race Through the Realm of Long Context with Large Language Models
von: Zhang, Lei, et al.
Veröffentlicht: (2023)