C-RADIOv4 (Tech Report)
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ranzinger, Mike, Heinrich, Greg, McCarthy, Collin, Kautz, Jan, Tao, Andrew, Catanzaro, Bryan, Molchanov, Pavlo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
von: Heinrich, Greg, et al.
Veröffentlicht: (2024)
von: Heinrich, Greg, et al.
Veröffentlicht: (2024)
FeatSharp: Your Vision Model Features, Sharper
von: Ranzinger, Mike, et al.
Veröffentlicht: (2025)
von: Ranzinger, Mike, et al.
Veröffentlicht: (2025)
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
von: Ranzinger, Mike, et al.
Veröffentlicht: (2023)
von: Ranzinger, Mike, et al.
Veröffentlicht: (2023)
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
von: Ranzinger, Mike, et al.
Veröffentlicht: (2024)
von: Ranzinger, Mike, et al.
Veröffentlicht: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
VILA: On Pre-training for Visual Language Models
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
LITA: Language Instructed Temporal-Localization Assistant
von: Huang, De-An, et al.
Veröffentlicht: (2024)
von: Huang, De-An, et al.
Veröffentlicht: (2024)
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
von: Li, Jiefeng, et al.
Veröffentlicht: (2024)
von: Li, Jiefeng, et al.
Veröffentlicht: (2024)
GSPN-2: Efficient Parallel Sequence Modeling
von: Wang, Hongjun, et al.
Veröffentlicht: (2025)
von: Wang, Hongjun, et al.
Veröffentlicht: (2025)
A Scalable Machine Learning Pipeline for Building Footprint Detection in Historical Maps
von: McCarthy, Annemarie
Veröffentlicht: (2025)
von: McCarthy, Annemarie
Veröffentlicht: (2025)
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
von: Jiang, Yitong, et al.
Veröffentlicht: (2026)
VILA$^2$: VILA Augmented VILA
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
von: Fang, Yunhao, et al.
Veröffentlicht: (2024)
X-VILA: Cross-Modality Alignment for Large Language Model
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
TwinTURBO: Semi-Supervised Fine-Tuning of Foundation Models via Mutual Information Decompositions for Downstream Task and Latent Spaces
von: Quétant, Guillaume, et al.
Veröffentlicht: (2025)
von: Quétant, Guillaume, et al.
Veröffentlicht: (2025)
Step Out and Seek Around: On Warm-Start Training with Incremental Data
von: Shen, Maying, et al.
Veröffentlicht: (2024)
von: Shen, Maying, et al.
Veröffentlicht: (2024)
ViR: Towards Efficient Vision Retention Backbones
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
OMCAT: Omni Context Aware Transformer
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
3D Aware Region Prompted Vision Language Model
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2025)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2025)
Adaptive Sharpness-Aware Pruning for Robust Sparse Networks
von: Bair, Anna, et al.
Veröffentlicht: (2023)
von: Bair, Anna, et al.
Veröffentlicht: (2023)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
MambaVision: A Hybrid Mamba-Transformer Vision Backbone
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2024)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2024)
Grounded 3D-Aware Spatial Vision-Language Modeling
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2026)
Advancing Weight and Channel Sparsification with Enhanced Saliency
von: Sun, Xinglong, et al.
Veröffentlicht: (2025)
von: Sun, Xinglong, et al.
Veröffentlicht: (2025)
neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing
von: Perrett, Toby, et al.
Veröffentlicht: (2026)
von: Perrett, Toby, et al.
Veröffentlicht: (2026)
FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
von: Liu, Hengyu, et al.
Veröffentlicht: (2025)
von: Liu, Hengyu, et al.
Veröffentlicht: (2025)
Scaling RL to Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
von: Chen, Yukang, et al.
Veröffentlicht: (2025)
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
von: Chen, Guo, et al.
Veröffentlicht: (2025)
von: Chen, Guo, et al.
Veröffentlicht: (2025)
MapVision: CVPR 2024 Autonomous Grand Challenge Mapless Driving Tech Report
von: Yang, Zhongyu, et al.
Veröffentlicht: (2024)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2024)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
DoRA: Weight-Decomposed Low-Rank Adaptation
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2024)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2024)
Perspective-Equivariant Fine-tuning for Multispectral Demosaicing without Ground Truth
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
Linked Adapters: Linking Past and Future to Present for Effective Continual Learning
von: Chandra, Dupati Srikar, et al.
Veröffentlicht: (2024)
von: Chandra, Dupati Srikar, et al.
Veröffentlicht: (2024)
Mixture of Gaussian-distributed Prototypes with Generative Modelling for Interpretable and Trustworthy Image Recognition
von: Wang, Chong, et al.
Veröffentlicht: (2023)
von: Wang, Chong, et al.
Veröffentlicht: (2023)
FoVA-Depth: Field-of-View Agnostic Depth Estimation for Cross-Dataset Generalization
von: Lichy, Daniel, et al.
Veröffentlicht: (2024)
von: Lichy, Daniel, et al.
Veröffentlicht: (2024)
AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion
von: Huang, Yangyi, et al.
Veröffentlicht: (2025)
von: Huang, Yangyi, et al.
Veröffentlicht: (2025)
nvTorchCam: An Open-source Library for Camera-Agnostic Differentiable Geometric Vision
von: Lichy, Daniel, et al.
Veröffentlicht: (2024)
von: Lichy, Daniel, et al.
Veröffentlicht: (2024)
A Lightweight Large Vision-language Model for Multimodal Medical Images
von: Alsinglawi, Belal, et al.
Veröffentlicht: (2025)
von: Alsinglawi, Belal, et al.
Veröffentlicht: (2025)
Hi-OSCAR: Hierarchical Open-set Classifier for Human Activity Recognition
von: McCarthy, Conor, et al.
Veröffentlicht: (2025)
von: McCarthy, Conor, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
von: Heinrich, Greg, et al.
Veröffentlicht: (2024) -
FeatSharp: Your Vision Model Features, Sharper
von: Ranzinger, Mike, et al.
Veröffentlicht: (2025) -
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
von: Ranzinger, Mike, et al.
Veröffentlicht: (2023) -
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
von: Ranzinger, Mike, et al.
Veröffentlicht: (2024) -
FasterViT: Fast Vision Transformers with Hierarchical Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)