Spider: Any-to-Many Multimodal LLM
Fuente:
arXiv
Salvato in:
| Autori principali: | Lai, Jinxiang, Zhang, Jie, Liu, Jun, Li, Jian, Lu, Xiaocheng, Guo, Song |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
di: Lai, Jinxiang, et al.
Pubblicazione: (2025)
di: Lai, Jinxiang, et al.
Pubblicazione: (2025)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
di: Gu, Bohai, et al.
Pubblicazione: (2026)
di: Gu, Bohai, et al.
Pubblicazione: (2026)
Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot Learning
di: Liu, Ziming, et al.
Pubblicazione: (2024)
di: Liu, Ziming, et al.
Pubblicazione: (2024)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
di: Li, Yanlin, et al.
Pubblicazione: (2026)
di: Li, Yanlin, et al.
Pubblicazione: (2026)
TokenPacker: Efficient Visual Projector for Multimodal LLM
di: Li, Wentong, et al.
Pubblicazione: (2024)
di: Li, Wentong, et al.
Pubblicazione: (2024)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
di: Guo, Hailong, et al.
Pubblicazione: (2025)
di: Guo, Hailong, et al.
Pubblicazione: (2025)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
di: Zhan, Jun, et al.
Pubblicazione: (2024)
di: Zhan, Jun, et al.
Pubblicazione: (2024)
VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model
di: Lai, Jinxiang, et al.
Pubblicazione: (2026)
di: Lai, Jinxiang, et al.
Pubblicazione: (2026)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
di: Feng, Yingying, et al.
Pubblicazione: (2025)
di: Feng, Yingying, et al.
Pubblicazione: (2025)
InstructSAM: Segment Any Instance with Any Instructions
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
di: Li, Bingchen, et al.
Pubblicazione: (2024)
di: Li, Bingchen, et al.
Pubblicazione: (2024)
Enhancing Multi-Class Anomaly Detection via Diffusion Refinement with Dual Conditioning
di: Zhan, Jiawei, et al.
Pubblicazione: (2024)
di: Zhan, Jiawei, et al.
Pubblicazione: (2024)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
di: Sun, Qichen, et al.
Pubblicazione: (2025)
di: Sun, Qichen, et al.
Pubblicazione: (2025)
FreeTuner: Any Subject in Any Style with Training-free Diffusion
di: Xu, Youcan, et al.
Pubblicazione: (2024)
di: Xu, Youcan, et al.
Pubblicazione: (2024)
DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning
di: Bai, Sikai, et al.
Pubblicazione: (2024)
di: Bai, Sikai, et al.
Pubblicazione: (2024)
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
di: Astruc, Guillaume, et al.
Pubblicazione: (2024)
di: Astruc, Guillaume, et al.
Pubblicazione: (2024)
X-Pose: Detecting Any Keypoints
di: Yang, Jie, et al.
Pubblicazione: (2023)
di: Yang, Jie, et al.
Pubblicazione: (2023)
Any Resolution Any Geometry: From Multi-View To Multi-Patch
di: Cui, Wenqing, et al.
Pubblicazione: (2026)
di: Cui, Wenqing, et al.
Pubblicazione: (2026)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
di: Lin, Xin, et al.
Pubblicazione: (2025)
di: Lin, Xin, et al.
Pubblicazione: (2025)
Anything in Any Scene: Photorealistic Video Object Insertion
di: Bai, Chen, et al.
Pubblicazione: (2024)
di: Bai, Chen, et al.
Pubblicazione: (2024)
LG-CAV: Train Any Concept Activation Vector with Language Guidance
di: Huang, Qihan, et al.
Pubblicazione: (2024)
di: Huang, Qihan, et al.
Pubblicazione: (2024)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
di: Huang, Muye, et al.
Pubblicazione: (2025)
di: Huang, Muye, et al.
Pubblicazione: (2025)
OutfitAnyone: Ultra-high Quality Virtual Try-On for Any Clothing and Any Person
di: Sun, Ke, et al.
Pubblicazione: (2024)
di: Sun, Ke, et al.
Pubblicazione: (2024)
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
di: Peng, Xinge, et al.
Pubblicazione: (2026)
di: Peng, Xinge, et al.
Pubblicazione: (2026)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
di: Li, Po-han, et al.
Pubblicazione: (2024)
di: Li, Po-han, et al.
Pubblicazione: (2024)
Animate Any Character in Any World
di: Wang, Yitong, et al.
Pubblicazione: (2025)
di: Wang, Yitong, et al.
Pubblicazione: (2025)
Dual Expert Distillation Network for Generalized Zero-Shot Learning
di: Rao, Zhijie, et al.
Pubblicazione: (2024)
di: Rao, Zhijie, et al.
Pubblicazione: (2024)
MatchDet: A Collaborative Framework for Image Matching and Object Detection
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
Clustered-patch Element Connection for Few-shot Learning
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
di: Lai, Jinxiang, et al.
Pubblicazione: (2023)
SpiderCam: Low-Power Snapshot Depth from Differential Defocus
di: Ferreira, Marcos A., et al.
Pubblicazione: (2026)
di: Ferreira, Marcos A., et al.
Pubblicazione: (2026)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
di: Wu, Changwei, et al.
Pubblicazione: (2025)
di: Wu, Changwei, et al.
Pubblicazione: (2025)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
di: Yeo, Juan, et al.
Pubblicazione: (2025)
di: Yeo, Juan, et al.
Pubblicazione: (2025)
AnyTSR: Any-Scale Thermal Super-Resolution for UAV
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
Motion Anything: Any to Motion Generation
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
di: Li, Shufan, et al.
Pubblicazione: (2023)
di: Li, Shufan, et al.
Pubblicazione: (2023)
Aligning Multimodal LLM with Human Preference: A Survey
di: Yu, Tao, et al.
Pubblicazione: (2025)
di: Yu, Tao, et al.
Pubblicazione: (2025)
AnyDepth: Depth Estimation Made Easy
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
UniDAC: Universal Metric Depth Estimation for Any Camera
di: Ganesan, Girish Chandar, et al.
Pubblicazione: (2026)
di: Ganesan, Girish Chandar, et al.
Pubblicazione: (2026)
tSF: Transformer-based Semantic Filter for Few-Shot Learning
di: Lai, Jinxiang, et al.
Pubblicazione: (2022)
di: Lai, Jinxiang, et al.
Pubblicazione: (2022)
Documenti analoghi
-
BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
di: Lai, Jinxiang, et al.
Pubblicazione: (2025) -
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
di: Gu, Bohai, et al.
Pubblicazione: (2026) -
Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot Learning
di: Liu, Ziming, et al.
Pubblicazione: (2024) -
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
di: Li, Yanlin, et al.
Pubblicazione: (2026) -
TokenPacker: Efficient Visual Projector for Multimodal LLM
di: Li, Wentong, et al.
Pubblicazione: (2024)