Saved in:
| Main Authors: | Zou, Bocheng, Cai, Mu, Stanley, Mark, Lu, Dingfu, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.25744 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuRF: Multi-Baseline Radiance Fields
by: Xu, Haofei, et al.
Published: (2023)
by: Xu, Haofei, et al.
Published: (2023)
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
by: Zou, Bocheng, et al.
Published: (2024)
by: Zou, Bocheng, et al.
Published: (2024)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
GS-Scale: Unlocking Large-Scale 3D Gaussian Splatting Training via Host Offloading
by: Lee, Donghyun, et al.
Published: (2025)
by: Lee, Donghyun, et al.
Published: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
by: Mantes, Albert Dominguez, et al.
Published: (2026)
by: Mantes, Albert Dominguez, et al.
Published: (2026)
Matryoshka Multimodal Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
by: Ming, Yuhang, et al.
Published: (2026)
by: Ming, Yuhang, et al.
Published: (2026)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
Enhancing Representation in Medical Vision-Language Foundation Models via Multi-Scale Information Extraction Techniques
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
by: Chae, Hyunsik, et al.
Published: (2025)
by: Chae, Hyunsik, et al.
Published: (2025)
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
by: Liu, Chenyang, et al.
Published: (2025)
by: Liu, Chenyang, et al.
Published: (2025)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
DivCon-NeRF: Diverse and Consistent Ray Augmentation for Few-Shot NeRF
by: Lee, Ingyun, et al.
Published: (2025)
by: Lee, Ingyun, et al.
Published: (2025)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
by: Xu, Muyu, et al.
Published: (2025)
by: Xu, Muyu, et al.
Published: (2025)
Agent Skills Should Go Beyond Text: The Case for Visual Skills
by: Xu, Binxiao, et al.
Published: (2026)
by: Xu, Binxiao, et al.
Published: (2026)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
by: Zhang, Jianrui, et al.
Published: (2024)
by: Zhang, Jianrui, et al.
Published: (2024)
MuST: Multi-Scale Transformers for Surgical Phase Recognition
by: Pérez, Alejandra, et al.
Published: (2024)
by: Pérez, Alejandra, et al.
Published: (2024)
MEIL-NeRF: Memory-Efficient Incremental Learning of Neural Radiance Fields
by: Chung, Jaeyoung, et al.
Published: (2022)
by: Chung, Jaeyoung, et al.
Published: (2022)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression Recognition
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Language-Guided Invariance Probing of Vision-Language Models
by: Lee, Jae Joong
Published: (2025)
by: Lee, Jae Joong
Published: (2025)
Unlocking the Potential of Operations Research for Multi-Graph Matching
by: Kahl, Max, et al.
Published: (2024)
by: Kahl, Max, et al.
Published: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
by: Miao, Yongzhu, et al.
Published: (2023)
by: Miao, Yongzhu, et al.
Published: (2023)
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
by: Zou, Yingshuang, et al.
Published: (2025)
by: Zou, Yingshuang, et al.
Published: (2025)
Active Prompt Learning in Vision Language Models
by: Bang, Jihwan, et al.
Published: (2023)
by: Bang, Jihwan, et al.
Published: (2023)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
ExBluRF: Efficient Radiance Fields for Extreme Motion Blurred Images
by: Lee, Dongwoo, et al.
Published: (2023)
by: Lee, Dongwoo, et al.
Published: (2023)
MuM: Multi-View Masked Image Modeling for 3D Vision
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
by: Zhang, Qizhe, et al.
Published: (2023)
by: Zhang, Qizhe, et al.
Published: (2023)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
by: Yu, Youngjoon, et al.
Published: (2024)
by: Yu, Youngjoon, et al.
Published: (2024)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
by: Jiang, Yitong, et al.
Published: (2026)
by: Jiang, Yitong, et al.
Published: (2026)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
by: Luo, Yulin, et al.
Published: (2024)
by: Luo, Yulin, et al.
Published: (2024)
Scaling Video Pretraining for Surgical Foundation Models
by: Lu, Sicheng, et al.
Published: (2026)
by: Lu, Sicheng, et al.
Published: (2026)
Similar Items
-
MuRF: Multi-Baseline Radiance Fields
by: Xu, Haofei, et al.
Published: (2023) -
VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation
by: Zou, Bocheng, et al.
Published: (2024) -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
by: Zeng, Ziyun, et al.
Published: (2026) -
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024) -
Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
by: Cai, Mu, et al.
Published: (2024)