Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Donggyun, Cho, Seongwoong, Kim, Semin, Luo, Chong, Hong, Seunghoon |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
par: Kim, Donggyun, et autres
Publié: (2025)
par: Kim, Donggyun, et autres
Publié: (2025)
Bridging the gap to real-world language-grounded visual concept learning
par: Jung, Whie, et autres
Publié: (2025)
par: Jung, Whie, et autres
Publié: (2025)
Universal Few-Shot Spatial Control for Diffusion Models
par: Nguyen, Kiet T., et autres
Publié: (2025)
par: Nguyen, Kiet T., et autres
Publié: (2025)
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
par: Cho, Yoorhim, et autres
Publié: (2025)
par: Cho, Yoorhim, et autres
Publié: (2025)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
par: Lee, Chanhyuk, et autres
Publié: (2025)
par: Lee, Chanhyuk, et autres
Publié: (2025)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
par: Cha, Yeonwoo, et autres
Publié: (2026)
par: Cha, Yeonwoo, et autres
Publié: (2026)
Meta-Controller: Few-Shot Imitation of Unseen Embodiments and Tasks in Continuous Control
par: Cho, Seongwoong, et autres
Publié: (2024)
par: Cho, Seongwoong, et autres
Publié: (2024)
Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
par: Lee, Dong-Guw, et autres
Publié: (2024)
par: Lee, Dong-Guw, et autres
Publié: (2024)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
par: Kim, Myunsoo, et autres
Publié: (2025)
par: Kim, Myunsoo, et autres
Publié: (2025)
Feature Augmentation based Test-Time Adaptation
par: Cho, Younggeol, et autres
Publié: (2024)
par: Cho, Younggeol, et autres
Publié: (2024)
MetaWeather: Few-Shot Weather-Degraded Image Restoration
par: Kim, Youngrae, et autres
Publié: (2023)
par: Kim, Youngrae, et autres
Publié: (2023)
Diffusion Model for Dense Matching
par: Nam, Jisu, et autres
Publié: (2023)
par: Nam, Jisu, et autres
Publié: (2023)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
par: Choi, Jiho, et autres
Publié: (2026)
par: Choi, Jiho, et autres
Publié: (2026)
Learning an Ensemble Token from Task-driven Priors in Facial Analysis
par: Seo, Sunyong, et autres
Publié: (2025)
par: Seo, Sunyong, et autres
Publié: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
par: Lee, Dong Hoon, et autres
Publié: (2024)
par: Lee, Dong Hoon, et autres
Publié: (2024)
Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing
par: Ko, Sukhun, et autres
Publié: (2026)
par: Ko, Sukhun, et autres
Publié: (2026)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
par: Jeong, Seunghoon, et autres
Publié: (2026)
par: Jeong, Seunghoon, et autres
Publié: (2026)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
par: Fan, Yue, et autres
Publié: (2024)
par: Fan, Yue, et autres
Publié: (2024)
From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios
par: Xia, Changliang, et autres
Publié: (2025)
par: Xia, Changliang, et autres
Publié: (2025)
Towards Motion-aware Referring Image Segmentation
par: Kim, Chaeyun, et autres
Publié: (2026)
par: Kim, Chaeyun, et autres
Publié: (2026)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
par: Hu, Yucheng, et autres
Publié: (2024)
par: Hu, Yucheng, et autres
Publié: (2024)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
par: Kim, Jaeyoon, et autres
Publié: (2025)
par: Kim, Jaeyoon, et autres
Publié: (2025)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
par: Hong, Sunghwan, et autres
Publié: (2024)
par: Hong, Sunghwan, et autres
Publié: (2024)
THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
par: Lee, Eunho, et autres
Publié: (2025)
par: Lee, Eunho, et autres
Publié: (2025)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
par: Gunawan, Agus, et autres
Publié: (2025)
par: Gunawan, Agus, et autres
Publié: (2025)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
par: Cho, Suhwan, et autres
Publié: (2024)
par: Cho, Suhwan, et autres
Publié: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
par: Cho, Hyeonwoo, et autres
Publié: (2026)
par: Cho, Hyeonwoo, et autres
Publié: (2026)
Data Augmentation For Small Object using Fast AutoAugment
par: Yoon, DaeEun, et autres
Publié: (2025)
par: Yoon, DaeEun, et autres
Publié: (2025)
Dual Prototype Attention for Unsupervised Video Object Segmentation
par: Cho, Suhwan, et autres
Publié: (2022)
par: Cho, Suhwan, et autres
Publié: (2022)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
par: Kim, Dong-Hee, et autres
Publié: (2024)
par: Kim, Dong-Hee, et autres
Publié: (2024)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
par: Kim, Junhyeok, et autres
Publié: (2025)
par: Kim, Junhyeok, et autres
Publié: (2025)
Efficient Prediction of Dense Visual Embeddings via Distillation and RGB-D Transformers
par: Fischedick, Söhnke Benedikt, et autres
Publié: (2026)
par: Fischedick, Söhnke Benedikt, et autres
Publié: (2026)
TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
par: Kim, Jeongyun, et autres
Publié: (2025)
par: Kim, Jeongyun, et autres
Publié: (2025)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
par: Zhao, Canyu, et autres
Publié: (2025)
par: Zhao, Canyu, et autres
Publié: (2025)
LRSLAM: Low-rank Representation of Signed Distance Fields in Dense Visual SLAM System
par: Park, Hongbeen, et autres
Publié: (2025)
par: Park, Hongbeen, et autres
Publié: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
par: Zhang, Dong, et autres
Publié: (2025)
par: Zhang, Dong, et autres
Publié: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
par: Lei, Zhenxin, et autres
Publié: (2025)
par: Lei, Zhenxin, et autres
Publié: (2025)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
par: Kang, Ju Yeon, et autres
Publié: (2025)
par: Kang, Ju Yeon, et autres
Publié: (2025)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
par: Kim, Jaeyoon, et autres
Publié: (2024)
par: Kim, Jaeyoon, et autres
Publié: (2024)
Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics
par: Cho, Woojin, et autres
Publié: (2024)
par: Cho, Woojin, et autres
Publié: (2024)
Documents similaires
-
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
par: Kim, Donggyun, et autres
Publié: (2025) -
Bridging the gap to real-world language-grounded visual concept learning
par: Jung, Whie, et autres
Publié: (2025) -
Universal Few-Shot Spatial Control for Diffusion Models
par: Nguyen, Kiet T., et autres
Publié: (2025) -
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
par: Cho, Yoorhim, et autres
Publié: (2025) -
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
par: Lee, Chanhyuk, et autres
Publié: (2025)