Bridging the gap to real-world language-grounded visual concept learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Jung, Whie, Kim, Semin, Kim, Junee, Hong, Seunghoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Disentangled Representation Learning via Modular Compositional Bias
di: Jung, Whie, et al.
Pubblicazione: (2025)
di: Jung, Whie, et al.
Pubblicazione: (2025)
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
di: Jung, Whie, et al.
Pubblicazione: (2024)
di: Jung, Whie, et al.
Pubblicazione: (2024)
Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
di: Kim, Donggyun, et al.
Pubblicazione: (2024)
di: Kim, Donggyun, et al.
Pubblicazione: (2024)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
di: Cha, Yeonwoo, et al.
Pubblicazione: (2026)
di: Cha, Yeonwoo, et al.
Pubblicazione: (2026)
Human-like compositional learning of visually-grounded concepts using synthetic environments
di: Lin, Zijun, et al.
Pubblicazione: (2025)
di: Lin, Zijun, et al.
Pubblicazione: (2025)
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
di: Choi, Jiho, et al.
Pubblicazione: (2026)
di: Choi, Jiho, et al.
Pubblicazione: (2026)
Learning an Ensemble Token from Task-driven Priors in Facial Analysis
di: Seo, Sunyong, et al.
Pubblicazione: (2025)
di: Seo, Sunyong, et al.
Pubblicazione: (2025)
RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
di: Cho, Yoorhim, et al.
Pubblicazione: (2025)
Feature Augmentation based Test-Time Adaptation
di: Cho, Younggeol, et al.
Pubblicazione: (2024)
di: Cho, Younggeol, et al.
Pubblicazione: (2024)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
di: Jeong, Seunghoon, et al.
Pubblicazione: (2026)
di: Jeong, Seunghoon, et al.
Pubblicazione: (2026)
Towards Motion-aware Referring Image Segmentation
di: Kim, Chaeyun, et al.
Pubblicazione: (2026)
di: Kim, Chaeyun, et al.
Pubblicazione: (2026)
Universal Few-Shot Spatial Control for Diffusion Models
di: Nguyen, Kiet T., et al.
Pubblicazione: (2025)
di: Nguyen, Kiet T., et al.
Pubblicazione: (2025)
MetaWeather: Few-Shot Weather-Degraded Image Restoration
di: Kim, Youngrae, et al.
Pubblicazione: (2023)
di: Kim, Youngrae, et al.
Pubblicazione: (2023)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
di: Nguezet, Paul Valery, et al.
Pubblicazione: (2026)
di: Nguezet, Paul Valery, et al.
Pubblicazione: (2026)
THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
di: Lee, Eunho, et al.
Pubblicazione: (2025)
di: Lee, Eunho, et al.
Pubblicazione: (2025)
Bridging the gap in FER: addressing age bias in deep learning
di: Gaya-Morey, F. Xavier, et al.
Pubblicazione: (2025)
di: Gaya-Morey, F. Xavier, et al.
Pubblicazione: (2025)
TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
di: Kim, Jeongyun, et al.
Pubblicazione: (2025)
di: Kim, Jeongyun, et al.
Pubblicazione: (2025)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
di: Chen, Yanyuan, et al.
Pubblicazione: (2025)
di: Chen, Yanyuan, et al.
Pubblicazione: (2025)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
di: Liu, Xiaohong, et al.
Pubblicazione: (2024)
di: Liu, Xiaohong, et al.
Pubblicazione: (2024)
Sparse-DeRF: Deblurred Neural Radiance Fields from Sparse View
di: Lee, Dogyoon, et al.
Pubblicazione: (2024)
di: Lee, Dogyoon, et al.
Pubblicazione: (2024)
Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery
di: Huynh, Andy V., et al.
Pubblicazione: (2024)
di: Huynh, Andy V., et al.
Pubblicazione: (2024)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
di: Jung, Junyoung, et al.
Pubblicazione: (2026)
di: Jung, Junyoung, et al.
Pubblicazione: (2026)
BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation
di: Kim, Beomjun, et al.
Pubblicazione: (2025)
di: Kim, Beomjun, et al.
Pubblicazione: (2025)
ORIDa: Object-centric Real-world Image Composition Dataset
di: Kim, Jinwoo, et al.
Pubblicazione: (2025)
di: Kim, Jinwoo, et al.
Pubblicazione: (2025)
SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
di: Kim, Jungho, et al.
Pubblicazione: (2026)
di: Kim, Jungho, et al.
Pubblicazione: (2026)
Controllable Long-term Motion Generation with Extended Joint Targets
di: Lee, Eunjong, et al.
Pubblicazione: (2025)
di: Lee, Eunjong, et al.
Pubblicazione: (2025)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
di: Li, Mingjian, et al.
Pubblicazione: (2024)
di: Li, Mingjian, et al.
Pubblicazione: (2024)
Data Augmentation For Small Object using Fast AutoAugment
di: Yoon, DaeEun, et al.
Pubblicazione: (2025)
di: Yoon, DaeEun, et al.
Pubblicazione: (2025)
Improving Unsupervised Video Object Segmentation via Fake Flow Generation
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
di: Cho, Suhwan, et al.
Pubblicazione: (2024)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
di: Hyung, Junha, et al.
Pubblicazione: (2024)
di: Hyung, Junha, et al.
Pubblicazione: (2024)
DySurface: Consistent 4D Surface Reconstruction via Bridging Explicit Gaussians and Implicit Functions
di: Kim, Minje, et al.
Pubblicazione: (2026)
di: Kim, Minje, et al.
Pubblicazione: (2026)
ContextMix: A context-aware data augmentation method for industrial visual inspection systems
di: Kim, Hyungmin, et al.
Pubblicazione: (2024)
di: Kim, Hyungmin, et al.
Pubblicazione: (2024)
Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
di: Bevandić, Petra, et al.
Pubblicazione: (2022)
di: Bevandić, Petra, et al.
Pubblicazione: (2022)
Dual Prototype Attention for Unsupervised Video Object Segmentation
di: Cho, Suhwan, et al.
Pubblicazione: (2022)
di: Cho, Suhwan, et al.
Pubblicazione: (2022)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
di: Lee, Sanghyeon, et al.
Pubblicazione: (2026)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2026)
Neuromorphic visual attention for Sign-language recognition on SpiNNaker
di: Liskova, Sarka, et al.
Pubblicazione: (2026)
di: Liskova, Sarka, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Disentangled Representation Learning via Modular Compositional Bias
di: Jung, Whie, et al.
Pubblicazione: (2025) -
Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
di: Jung, Whie, et al.
Pubblicazione: (2024) -
Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
di: Kim, Donggyun, et al.
Pubblicazione: (2024) -
Training-Free Refinement of Flow Matching with Divergence-based Sampling
di: Cha, Yeonwoo, et al.
Pubblicazione: (2026) -
Human-like compositional learning of visually-grounded concepts using synthetic environments
di: Lin, Zijun, et al.
Pubblicazione: (2025)