Gespeichert in:
| Hauptverfasser: | Liu, Hong, Lu, Yitong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2411.16236 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interfacing Foundation Models' Embeddings
von: Zou, Xueyan, et al.
Veröffentlicht: (2023)
von: Zou, Xueyan, et al.
Veröffentlicht: (2023)
Pixel Sentence Representation Learning
von: Xiao, Chenghao, et al.
Veröffentlicht: (2024)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2024)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
Improving Arabic Multi-Label Emotion Classification using Stacked Embeddings and Hybrid Loss Function
von: Aslam, Muhammad Azeem, et al.
Veröffentlicht: (2024)
von: Aslam, Muhammad Azeem, et al.
Veröffentlicht: (2024)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
von: Tian, Junjiao, et al.
Veröffentlicht: (2024)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Docopilot: Improving Multimodal Models for Document-Level Understanding
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
von: Li, Kaican, et al.
Veröffentlicht: (2025)
von: Li, Kaican, et al.
Veröffentlicht: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
von: Wang, Chenguang, et al.
Veröffentlicht: (2024)
von: Wang, Chenguang, et al.
Veröffentlicht: (2024)
MiniMax-01: Scaling Foundation Models with Lightning Attention
von: MiniMax, et al.
Veröffentlicht: (2025)
von: MiniMax, et al.
Veröffentlicht: (2025)
Historical Test-time Prompt Tuning for Vision Foundation Models
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2024)
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
von: Ma, Guoqing, et al.
Veröffentlicht: (2025)
von: Ma, Guoqing, et al.
Veröffentlicht: (2025)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
MiMo-Embodied: X-Embodied Foundation Model Technical Report
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models
von: Diesendruck, Maurice, et al.
Veröffentlicht: (2024)
von: Diesendruck, Maurice, et al.
Veröffentlicht: (2024)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
von: Xu, Xiaohao, et al.
Veröffentlicht: (2024)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
von: Yang, Yuhao, et al.
Veröffentlicht: (2025)
von: Yang, Yuhao, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucination via Robust Local Perception Search
von: Gao, Zixian, et al.
Veröffentlicht: (2025)
von: Gao, Zixian, et al.
Veröffentlicht: (2025)
EPEE: Towards Efficient and Effective Foundation Models in Biomedicine
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
Measuring and Improving Persuasiveness of Large Language Models
von: Singh, Somesh, et al.
Veröffentlicht: (2024)
von: Singh, Somesh, et al.
Veröffentlicht: (2024)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
von: Chen, Zhihong, et al.
Veröffentlicht: (2024)
von: Chen, Zhihong, et al.
Veröffentlicht: (2024)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
MIEB: Massive Image Embedding Benchmark
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
von: Yin, Wei, et al.
Veröffentlicht: (2022)
von: Yin, Wei, et al.
Veröffentlicht: (2022)
Dynamic Token Reweighting for Robust Vision-Language Models
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2025)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2025)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interfacing Foundation Models' Embeddings
von: Zou, Xueyan, et al.
Veröffentlicht: (2023) -
Pixel Sentence Representation Learning
von: Xiao, Chenghao, et al.
Veröffentlicht: (2024) -
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024) -
Improving Arabic Multi-Label Emotion Classification using Stacked Embeddings and Hybrid Loss Function
von: Aslam, Muhammad Azeem, et al.
Veröffentlicht: (2024) -
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
von: Wang, Yeyuan, et al.
Veröffentlicht: (2025)