4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Bachmann, Roman, Kar, Oğuzhan Fatih, Mizrahi, David, Garjani, Ali, Gao, Mingfei, Griffiths, David, Hu, Jiaming, Dehghan, Afshin, Zamir, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
by: Bachmann, Roman, et al.
Published: (2025)
by: Bachmann, Roman, et al.
Published: (2025)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
(1D) Ordered Tokens Enable Efficient Test-Time Search
by: Gao, Zhitong, et al.
Published: (2026)
by: Gao, Zhitong, et al.
Published: (2026)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Unraveling the Key Components of OOD Generalization via Diversification
by: Benoit, Harold, et al.
Published: (2023)
by: Benoit, Harold, et al.
Published: (2023)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
by: Guo, Hailong, et al.
Published: (2025)
by: Guo, Hailong, et al.
Published: (2025)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
by: Feng, Yingying, et al.
Published: (2025)
by: Feng, Yingying, et al.
Published: (2025)
Symbolic Representation for Any-to-Any Generative Tasks
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
Cubify Anything: Scaling Indoor 3D Object Detection
by: Lazarow, Justin, et al.
Published: (2024)
by: Lazarow, Justin, et al.
Published: (2024)
Segment Any Medical Model Extended
by: Liu, Yihao, et al.
Published: (2024)
by: Liu, Yihao, et al.
Published: (2024)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
by: Chen, Haoyang, et al.
Published: (2026)
by: Chen, Haoyang, et al.
Published: (2026)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling
by: Seong, Kiyoung, et al.
Published: (2026)
by: Seong, Kiyoung, et al.
Published: (2026)
Tracking and Segmenting Anything in Any Modality
by: Zhang, Tianlu, et al.
Published: (2025)
by: Zhang, Tianlu, et al.
Published: (2025)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
by: Wu, Changwei, et al.
Published: (2025)
by: Wu, Changwei, et al.
Published: (2025)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
by: Tang, Yiwen, et al.
Published: (2024)
by: Tang, Yiwen, et al.
Published: (2024)
SAMM (Segment Any Medical Model): A 3D Slicer Integration to SAM
by: Liu, Yihao, et al.
Published: (2023)
by: Liu, Yihao, et al.
Published: (2023)
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
by: Salehi, Sogand, et al.
Published: (2024)
by: Salehi, Sogand, et al.
Published: (2024)
Any Model, Any Place, Any Time: Get Remote Sensing Foundation Model Embeddings On Demand
by: Ye, Dingqi, et al.
Published: (2026)
by: Ye, Dingqi, et al.
Published: (2026)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
Segment Any Mesh
by: Tang, George, et al.
Published: (2024)
by: Tang, George, et al.
Published: (2024)
Symbol Emergence and The Solutions to Any Task
by: Bennett, Michael Timothy
Published: (2021)
by: Bennett, Michael Timothy
Published: (2021)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
Omni-Mol: Multitask Molecular Model for Any-to-any Modalities
by: Hu, Chengxin, et al.
Published: (2025)
by: Hu, Chengxin, et al.
Published: (2025)
Multi-Vector Index Compression in Any Modality
by: Qin, Hanxiang, et al.
Published: (2026)
by: Qin, Hanxiang, et al.
Published: (2026)
Structurally Prune Anything: Any Architecture, Any Framework, Any Time
by: Wang, Xun, et al.
Published: (2024)
by: Wang, Xun, et al.
Published: (2024)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Language Models Improve When Pretraining Data Matches Target Tasks
by: Mizrahi, David, et al.
Published: (2025)
by: Mizrahi, David, et al.
Published: (2025)
Animate Any Character in Any World
by: Wang, Yitong, et al.
Published: (2025)
by: Wang, Yitong, et al.
Published: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
by: Zhan, Wengyi, et al.
Published: (2024)
by: Zhan, Wengyi, et al.
Published: (2024)
SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality
by: Lei, Chenyang, et al.
Published: (2024)
by: Lei, Chenyang, et al.
Published: (2024)
Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
AnyPcc: Compressing Any Point Cloud with a Single Universal Model
by: Wang, Kangli, et al.
Published: (2025)
by: Wang, Kangli, et al.
Published: (2025)
NExT-GPT: Any-to-Any Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2023)
by: Wu, Shengqiong, et al.
Published: (2023)
Similar Items
-
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025) -
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
by: Bachmann, Roman, et al.
Published: (2025) -
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026) -
(1D) Ordered Tokens Enable Efficient Test-Time Search
by: Gao, Zhitong, et al.
Published: (2026) -
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)