TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Haokun, Wang, Teng, Ge, Yixiao, Ge, Yuying, Lu, Zhichao, Wei, Ying, Zhang, Qingfu, Sun, Zhenan, Shan, Ying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
von: Li, Bohao, et al.
Veröffentlicht: (2024)
von: Li, Bohao, et al.
Veröffentlicht: (2024)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
von: Ma, Shijie, et al.
Veröffentlicht: (2025)
von: Ma, Shijie, et al.
Veröffentlicht: (2025)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
SEED-Story: Multimodal Long Story Generation with Large Language Model
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
von: Pu, Junfu, et al.
Veröffentlicht: (2025)
von: Pu, Junfu, et al.
Veröffentlicht: (2025)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
Supervised Fine-tuning in turn Improves Visual Foundation Models
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Aligning Latent Spaces with Flow Priors
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
Scalable Image Tokenization with Index Backpropagation Quantization
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Luo, Zhuoyan, et al.
Veröffentlicht: (2024)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
ST-LLM: Large Language Models Are Effective Temporal Learners
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
SCHNet: SAM Marries CLIP for Human Parsing
von: Liu, Kunliang, et al.
Veröffentlicht: (2025)
von: Liu, Kunliang, et al.
Veröffentlicht: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
von: Ji, Yatai, et al.
Veröffentlicht: (2025)
von: Ji, Yatai, et al.
Veröffentlicht: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Evolving Interdependent Operators with Large Language Models for Multi-Objective Combinatorial Optimization
von: Qiu, Junhao, et al.
Veröffentlicht: (2026)
von: Qiu, Junhao, et al.
Veröffentlicht: (2026)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
YOLO-World: Real-Time Open-Vocabulary Object Detection
von: Cheng, Tianheng, et al.
Veröffentlicht: (2024)
von: Cheng, Tianheng, et al.
Veröffentlicht: (2024)
DINO-Tok: Adapting DINO for Visual Tokenizers
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
von: Cheng, Cheng, et al.
Veröffentlicht: (2023)
von: Cheng, Cheng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024) -
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
von: Li, Bohao, et al.
Veröffentlicht: (2024) -
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024) -
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025) -
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)