Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Donggyun, Cho, Seongwoong, Kim, Semin, Luo, Chong, Hong, Seunghoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
by: Kim, Donggyun, et al.
Published: (2025)
by: Kim, Donggyun, et al.
Published: (2025)
Bridging the gap to real-world language-grounded visual concept learning
by: Jung, Whie, et al.
Published: (2025)
by: Jung, Whie, et al.
Published: (2025)
Universal Few-Shot Spatial Control for Diffusion Models
by: Nguyen, Kiet T., et al.
Published: (2025)
by: Nguyen, Kiet T., et al.
Published: (2025)
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
by: Cho, Yoorhim, et al.
Published: (2025)
by: Cho, Yoorhim, et al.
Published: (2025)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
by: Lee, Chanhyuk, et al.
Published: (2025)
by: Lee, Chanhyuk, et al.
Published: (2025)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
by: Cha, Yeonwoo, et al.
Published: (2026)
by: Cha, Yeonwoo, et al.
Published: (2026)
Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control
by: Cho, Seongwoong, et al.
Published: (2024)
by: Cho, Seongwoong, et al.
Published: (2024)
Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
by: Lee, Dong-Guw, et al.
Published: (2024)
by: Lee, Dong-Guw, et al.
Published: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
Feature Augmentation based Test-Time Adaptation
by: Cho, Younggeol, et al.
Published: (2024)
by: Cho, Younggeol, et al.
Published: (2024)
MetaWeather: Few-Shot Weather-Degraded Image Restoration
by: Kim, Youngrae, et al.
Published: (2023)
by: Kim, Youngrae, et al.
Published: (2023)
Diffusion Model for Dense Matching
by: Nam, Jisu, et al.
Published: (2023)
by: Nam, Jisu, et al.
Published: (2023)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
by: Choi, Jiho, et al.
Published: (2026)
by: Choi, Jiho, et al.
Published: (2026)
Learning an Ensemble Token from Task-driven Priors in Facial Analysis
by: Seo, Sunyong, et al.
Published: (2025)
by: Seo, Sunyong, et al.
Published: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
by: Lee, Dong Hoon, et al.
Published: (2024)
by: Lee, Dong Hoon, et al.
Published: (2024)
Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
by: Ko, Sukhun, et al.
Published: (2026)
by: Ko, Sukhun, et al.
Published: (2026)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
by: Jeong, Seunghoon, et al.
Published: (2026)
by: Jeong, Seunghoon, et al.
Published: (2026)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios
by: Xia, Changliang, et al.
Published: (2025)
by: Xia, Changliang, et al.
Published: (2025)
Towards Motion-aware Referring Image Segmentation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
by: Kim, Jaeyoon, et al.
Published: (2025)
by: Kim, Jaeyoon, et al.
Published: (2025)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
by: Hong, Sunghwan, et al.
Published: (2024)
by: Hong, Sunghwan, et al.
Published: (2024)
THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
by: Lee, Eunho, et al.
Published: (2025)
by: Lee, Eunho, et al.
Published: (2025)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
Data Augmentation For Small Object using Fast AutoAugment
by: Yoon, DaeEun, et al.
Published: (2025)
by: Yoon, DaeEun, et al.
Published: (2025)
Dual Prototype Attention for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2022)
by: Cho, Suhwan, et al.
Published: (2022)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
by: Kim, Dong-Hee, et al.
Published: (2024)
by: Kim, Dong-Hee, et al.
Published: (2024)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
Efficient Prediction of Dense Visual Embeddings via Distillation and RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2026)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2026)
TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
by: Kim, Jeongyun, et al.
Published: (2025)
by: Kim, Jeongyun, et al.
Published: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
by: Park, Hongbeen, et al.
Published: (2025)
by: Park, Hongbeen, et al.
Published: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
by: Kim, Jaeyoon, et al.
Published: (2024)
by: Kim, Jaeyoon, et al.
Published: (2024)
Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics
by: Cho, Woojin, et al.
Published: (2024)
by: Cho, Woojin, et al.
Published: (2024)
Similar Items
-
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
by: Kim, Donggyun, et al.
Published: (2025) -
Bridging the gap to real-world language-grounded visual concept learning
by: Jung, Whie, et al.
Published: (2025) -
Universal Few-Shot Spatial Control for Diffusion Models
by: Nguyen, Kiet T., et al.
Published: (2025) -
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
by: Cho, Yoorhim, et al.
Published: (2025) -
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
by: Lee, Chanhyuk, et al.
Published: (2025)