MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Bocheng, Cai, Mu, Stanley, Mark, Lu, Dingfu, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MuRF: Multi-Baseline Radiance Fields
von: Xu, Haofei, et al.
Veröffentlicht: (2023)
von: Xu, Haofei, et al.
Veröffentlicht: (2023)
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
von: Zou, Bocheng, et al.
Veröffentlicht: (2024)
von: Zou, Bocheng, et al.
Veröffentlicht: (2024)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
von: Wu, Yuqun, et al.
Veröffentlicht: (2024)
von: Wu, Yuqun, et al.
Veröffentlicht: (2024)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
GS-Scale: Unlocking Large-Scale 3D Gaussian Splatting Training via Host Offloading
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)
Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
von: Mantes, Albert Dominguez, et al.
Veröffentlicht: (2026)
von: Mantes, Albert Dominguez, et al.
Veröffentlicht: (2026)
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
von: Ming, Yuhang, et al.
Veröffentlicht: (2026)
von: Ming, Yuhang, et al.
Veröffentlicht: (2026)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
von: Huang, Weijian, et al.
Veröffentlicht: (2024)
von: Huang, Weijian, et al.
Veröffentlicht: (2024)
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
von: Liu, Chenyang, et al.
Veröffentlicht: (2025)
von: Liu, Chenyang, et al.
Veröffentlicht: (2025)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
von: Xu, Muyu, et al.
Veröffentlicht: (2025)
von: Xu, Muyu, et al.
Veröffentlicht: (2025)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
DivCon-NeRF: Diverse and Consistent Ray Augmentation for Few-Shot NeRF
von: Lee, Ingyun, et al.
Veröffentlicht: (2025)
von: Lee, Ingyun, et al.
Veröffentlicht: (2025)
Matryoshka Multimodal Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
von: Shang, Yuzhang, et al.
Veröffentlicht: (2024)
Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
MuST: Multi-Scale Transformers for Surgical Phase Recognition
von: Pérez, Alejandra, et al.
Veröffentlicht: (2024)
von: Pérez, Alejandra, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Operations Research for Multi-Graph Matching
von: Kahl, Max, et al.
Veröffentlicht: (2024)
von: Kahl, Max, et al.
Veröffentlicht: (2024)
MEIL-NeRF: Memory-Efficient Incremental Learning of Neural Radiance Fields
von: Chung, Jaeyoung, et al.
Veröffentlicht: (2022)
von: Chung, Jaeyoung, et al.
Veröffentlicht: (2022)
Agent Skills Should Go Beyond Text: The Case for Visual Skills
von: Xu, Binxiao, et al.
Veröffentlicht: (2026)
von: Xu, Binxiao, et al.
Veröffentlicht: (2026)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
von: Miao, Yongzhu, et al.
Veröffentlicht: (2023)
von: Miao, Yongzhu, et al.
Veröffentlicht: (2023)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
von: Li, Nanxi, et al.
Veröffentlicht: (2026)
von: Li, Nanxi, et al.
Veröffentlicht: (2026)
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
Active Prompt Learning in Vision Language Models
von: Bang, Jihwan, et al.
Veröffentlicht: (2023)
von: Bang, Jihwan, et al.
Veröffentlicht: (2023)
Language-Guided Invariance Probing of Vision-Language Models
von: Lee, Jae Joong
Veröffentlicht: (2025)
von: Lee, Jae Joong
Veröffentlicht: (2025)
ExBluRF: Efficient Radiance Fields for Extreme Motion Blurred Images
von: Lee, Dongwoo, et al.
Veröffentlicht: (2023)
von: Lee, Dongwoo, et al.
Veröffentlicht: (2023)
Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
von: Yu, Youngjoon, et al.
Veröffentlicht: (2024)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
von: Hur, Jiwan, et al.
Veröffentlicht: (2024)
von: Hur, Jiwan, et al.
Veröffentlicht: (2024)
Scaling Video Pretraining for Surgical Foundation Models
von: Lu, Sicheng, et al.
Veröffentlicht: (2026)
von: Lu, Sicheng, et al.
Veröffentlicht: (2026)
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
von: He, Xunjie, et al.
Veröffentlicht: (2025)
von: He, Xunjie, et al.
Veröffentlicht: (2025)
OV-NeRF: Open-vocabulary Neural Radiance Fields with Vision and Language Foundation Models for 3D Semantic Understanding
von: Liao, Guibiao, et al.
Veröffentlicht: (2024)
von: Liao, Guibiao, et al.
Veröffentlicht: (2024)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation
von: Zhang, Haojie, et al.
Veröffentlicht: (2026)
von: Zhang, Haojie, et al.
Veröffentlicht: (2026)
MuM: Multi-View Masked Image Modeling for 3D Vision
von: Nordström, David, et al.
Veröffentlicht: (2025)
von: Nordström, David, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MuRF: Multi-Baseline Radiance Fields
von: Xu, Haofei, et al.
Veröffentlicht: (2023) -
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
von: Zou, Bocheng, et al.
Veröffentlicht: (2024) -
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
von: Wu, Yuqun, et al.
Veröffentlicht: (2024) -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026) -
GS-Scale: Unlocking Large-Scale 3D Gaussian Splatting Training via Host Offloading
von: Lee, Donghyun, et al.
Veröffentlicht: (2025)