Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Donggyun, Cho, Seongwoong, Kim, Semin, Luo, Chong, Hong, Seunghoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
Bridging the gap to real-world language-grounded visual concept learning
di: Jung, Whie, et al.
Pubblicazione: (2025)
di: Jung, Whie, et al.
Pubblicazione: (2025)
Universal Few-Shot Spatial Control for Diffusion Models
di: Nguyen, Kiet T., et al.
Pubblicazione: (2025)
di: Nguyen, Kiet T., et al.
Pubblicazione: (2025)
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
di: Cha, Yeonwoo, et al.
Pubblicazione: (2026)
di: Cha, Yeonwoo, et al.
Pubblicazione: (2026)
Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control
di: Cho, Seongwoong, et al.
Pubblicazione: (2024)
di: Cho, Seongwoong, et al.
Pubblicazione: (2024)
Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
di: Lee, Dong-Guw, et al.
Pubblicazione: (2024)
di: Lee, Dong-Guw, et al.
Pubblicazione: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
di: Kim, Myunsoo, et al.
Pubblicazione: (2025)
di: Kim, Myunsoo, et al.
Pubblicazione: (2025)
Feature Augmentation based Test-Time Adaptation
di: Cho, Younggeol, et al.
Pubblicazione: (2024)
di: Cho, Younggeol, et al.
Pubblicazione: (2024)
MetaWeather: Few-Shot Weather-Degraded Image Restoration
di: Kim, Youngrae, et al.
Pubblicazione: (2023)
di: Kim, Youngrae, et al.
Pubblicazione: (2023)
Diffusion Model for Dense Matching
di: Nam, Jisu, et al.
Pubblicazione: (2023)
di: Nam, Jisu, et al.
Pubblicazione: (2023)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
di: Choi, Jiho, et al.
Pubblicazione: (2026)
di: Choi, Jiho, et al.
Pubblicazione: (2026)
Learning an Ensemble Token from Task-driven Priors in Facial Analysis
di: Seo, Sunyong, et al.
Pubblicazione: (2025)
di: Seo, Sunyong, et al.
Pubblicazione: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
di: Ko, Sukhun, et al.
Pubblicazione: (2026)
di: Ko, Sukhun, et al.
Pubblicazione: (2026)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
di: Jeong, Seunghoon, et al.
Pubblicazione: (2026)
di: Jeong, Seunghoon, et al.
Pubblicazione: (2026)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios
di: Xia, Changliang, et al.
Pubblicazione: (2025)
di: Xia, Changliang, et al.
Pubblicazione: (2025)
Towards Motion-aware Referring Image Segmentation
di: Kim, Chaeyun, et al.
Pubblicazione: (2026)
di: Kim, Chaeyun, et al.
Pubblicazione: (2026)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
di: Hu, Yucheng, et al.
Pubblicazione: (2024)
di: Hu, Yucheng, et al.
Pubblicazione: (2024)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
di: Kim, Jaeyoon, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoon, et al.
Pubblicazione: (2025)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
di: Lee, Eunho, et al.
Pubblicazione: (2025)
di: Lee, Eunho, et al.
Pubblicazione: (2025)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
di: Gunawan, Agus, et al.
Pubblicazione: (2025)
di: Gunawan, Agus, et al.
Pubblicazione: (2025)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
di: Cho, Hyeonwoo, et al.
Pubblicazione: (2026)
di: Cho, Hyeonwoo, et al.
Pubblicazione: (2026)
Data Augmentation For Small Object using Fast AutoAugment
di: Yoon, DaeEun, et al.
Pubblicazione: (2025)
di: Yoon, DaeEun, et al.
Pubblicazione: (2025)
Dual Prototype Attention for Unsupervised Video Object Segmentation
di: Cho, Suhwan, et al.
Pubblicazione: (2022)
di: Cho, Suhwan, et al.
Pubblicazione: (2022)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
di: Kim, Dong-Hee, et al.
Pubblicazione: (2024)
di: Kim, Dong-Hee, et al.
Pubblicazione: (2024)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
di: Kim, Junhyeok, et al.
Pubblicazione: (2025)
di: Kim, Junhyeok, et al.
Pubblicazione: (2025)
Efficient Prediction of Dense Visual Embeddings via Distillation and RGB-D Transformers
di: Fischedick, Söhnke Benedikt, et al.
Pubblicazione: (2026)
di: Fischedick, Söhnke Benedikt, et al.
Pubblicazione: (2026)
TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
di: Kim, Jeongyun, et al.
Pubblicazione: (2025)
di: Kim, Jeongyun, et al.
Pubblicazione: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
di: Park, Hongbeen, et al.
Pubblicazione: (2025)
di: Park, Hongbeen, et al.
Pubblicazione: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
di: Zhang, Dong, et al.
Pubblicazione: (2025)
di: Zhang, Dong, et al.
Pubblicazione: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
di: Kim, Jaeyoon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyoon, et al.
Pubblicazione: (2024)
Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics
di: Cho, Woojin, et al.
Pubblicazione: (2024)
di: Cho, Woojin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
di: Kim, Donggyun, et al.
Pubblicazione: (2025) -
Bridging the gap to real-world language-grounded visual concept learning
di: Jung, Whie, et al.
Pubblicazione: (2025) -
Universal Few-Shot Spatial Control for Diffusion Models
di: Nguyen, Kiet T., et al.
Pubblicazione: (2025) -
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
di: Cho, Yoorhim, et al.
Pubblicazione: (2025) -
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)