"Principal Components" Enable A New Language of Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Wen, Xin, Zhao, Bingchen, Elezi, Ismail, Deng, Jiankang, Qi, Xiaojuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
G3DR: Generative 3D Reconstruction in ImageNet
di: Reddy, Pradyumna, et al.
Pubblicazione: (2024)
di: Reddy, Pradyumna, et al.
Pubblicazione: (2024)
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
di: Miles, Roy, et al.
Pubblicazione: (2024)
di: Miles, Roy, et al.
Pubblicazione: (2024)
Deep Active Learning: A Reality Check
di: Gashi, Edrina, et al.
Pubblicazione: (2024)
di: Gashi, Edrina, et al.
Pubblicazione: (2024)
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2025)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2025)
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
di: Miles, Roy, et al.
Pubblicazione: (2024)
di: Miles, Roy, et al.
Pubblicazione: (2024)
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
di: Ma, Chengcheng, et al.
Pubblicazione: (2023)
di: Ma, Chengcheng, et al.
Pubblicazione: (2023)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
di: Toker, Aysim, et al.
Pubblicazione: (2025)
di: Toker, Aysim, et al.
Pubblicazione: (2025)
Fractal Calibration for long-tailed object detection
di: Alexandridis, Konstantinos Panagiotis, et al.
Pubblicazione: (2024)
di: Alexandridis, Konstantinos Panagiotis, et al.
Pubblicazione: (2024)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
di: Wen, Xin, et al.
Pubblicazione: (2025)
di: Wen, Xin, et al.
Pubblicazione: (2025)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
di: Choi, Yura, et al.
Pubblicazione: (2026)
di: Choi, Yura, et al.
Pubblicazione: (2026)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
di: Wen, Xin, et al.
Pubblicazione: (2024)
di: Wen, Xin, et al.
Pubblicazione: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
di: Xie, Yin, et al.
Pubblicazione: (2024)
di: Xie, Yin, et al.
Pubblicazione: (2024)
Region-based Cluster Discrimination for Visual Representation Learning
di: Xie, Yin, et al.
Pubblicazione: (2025)
di: Xie, Yin, et al.
Pubblicazione: (2025)
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
di: Khan, Mohammad Sadil, et al.
Pubblicazione: (2026)
di: Khan, Mohammad Sadil, et al.
Pubblicazione: (2026)
Interpretable Text-Guided Image Clustering via Iterative Search
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
di: Zhao, Shizhen, et al.
Pubblicazione: (2025)
di: Zhao, Shizhen, et al.
Pubblicazione: (2025)
Vision Foundation Models as Generalist Tokenizers for Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2026)
di: Zheng, Anlin, et al.
Pubblicazione: (2026)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
di: Zuo, Ronglai, et al.
Pubblicazione: (2026)
di: Zuo, Ronglai, et al.
Pubblicazione: (2026)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
di: Li, Bingchen, et al.
Pubblicazione: (2024)
di: Li, Bingchen, et al.
Pubblicazione: (2024)
Generalized Category Discovery under the Long-Tailed Distribution
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
di: Zhao, Bingchen, et al.
Pubblicazione: (2025)
Learning from Neighbors: Category Extrapolation for Long-Tail Learning
di: Zhao, Shizhen, et al.
Pubblicazione: (2024)
di: Zhao, Shizhen, et al.
Pubblicazione: (2024)
Hyperspectral Image Spectral-Spatial Feature Extraction via Tensor Principal Component Analysis
di: Ren, Yuemei, et al.
Pubblicazione: (2024)
di: Ren, Yuemei, et al.
Pubblicazione: (2024)
MambaCSR: Dual-Interleaved Scanning for Compressed Image Super-Resolution With SSMs
di: Ren, Yulin, et al.
Pubblicazione: (2024)
di: Ren, Yulin, et al.
Pubblicazione: (2024)
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
di: Wang, Xijun, et al.
Pubblicazione: (2025)
di: Wang, Xijun, et al.
Pubblicazione: (2025)
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models
di: Zhang, Letian, et al.
Pubblicazione: (2023)
di: Zhang, Letian, et al.
Pubblicazione: (2023)
Eigenpatches -- Adversarial Patches from Principal Components
di: Bayer, Jens, et al.
Pubblicazione: (2023)
di: Bayer, Jens, et al.
Pubblicazione: (2023)
Feature Aligning Few shot Learning Method Using Local Descriptors Weighted Rules
di: Yan, Bingchen
Pubblicazione: (2024)
di: Yan, Bingchen
Pubblicazione: (2024)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
di: Cui, Jiequan, et al.
Pubblicazione: (2024)
di: Cui, Jiequan, et al.
Pubblicazione: (2024)
Can OOD Object Detectors Learn from Foundation Models?
di: Liu, Jiahui, et al.
Pubblicazione: (2024)
di: Liu, Jiahui, et al.
Pubblicazione: (2024)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
di: Cui, Siying, et al.
Pubblicazione: (2024)
di: Cui, Siying, et al.
Pubblicazione: (2024)
Robust Principal Component Analysis via Discriminant Sample Weight Learning
di: Deng, Yingzhuo, et al.
Pubblicazione: (2024)
di: Deng, Yingzhuo, et al.
Pubblicazione: (2024)
Unleashing Vision-Language Semantics for Deepfake Video Detection
di: Zhu, Jiawen, et al.
Pubblicazione: (2026)
di: Zhu, Jiawen, et al.
Pubblicazione: (2026)
Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation
di: Du, Zhipeng, et al.
Pubblicazione: (2023)
di: Du, Zhipeng, et al.
Pubblicazione: (2023)
MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
di: Ye, Xinyan, et al.
Pubblicazione: (2026)
di: Ye, Xinyan, et al.
Pubblicazione: (2026)
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery
di: Miao, Yunqi, et al.
Pubblicazione: (2024)
di: Miao, Yunqi, et al.
Pubblicazione: (2024)
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
di: Lu, Yanzuo, et al.
Pubblicazione: (2026)
di: Lu, Yanzuo, et al.
Pubblicazione: (2026)
Robust Principal Component Completion
di: Wang, Yinjian, et al.
Pubblicazione: (2026)
di: Wang, Yinjian, et al.
Pubblicazione: (2026)
Can 3D Vision-Language Models Truly Understand Natural Language?
di: Deng, Weipeng, et al.
Pubblicazione: (2024)
di: Deng, Weipeng, et al.
Pubblicazione: (2024)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
di: Song, Nan, et al.
Pubblicazione: (2025)
di: Song, Nan, et al.
Pubblicazione: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
di: Zhao, Bingchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
G3DR: Generative 3D Reconstruction in ImageNet
di: Reddy, Pradyumna, et al.
Pubblicazione: (2024) -
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
di: Miles, Roy, et al.
Pubblicazione: (2024) -
Deep Active Learning: A Reality Check
di: Gashi, Edrina, et al.
Pubblicazione: (2024) -
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2025) -
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
di: Miles, Roy, et al.
Pubblicazione: (2024)