Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jaeill, Hwang, Duhun, Lee, Eunjung, Suh, Jangwon, Kim, Jimyeong, Rhee, Wonjong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Feature Attribution Methods for Electrocardiogram
por: Suh, Jangwon, et al.
Publicado: (2022)
por: Suh, Jangwon, et al.
Publicado: (2022)
Towards a Better Evaluation of Out-of-Domain Generalization
por: Hwang, Duhun, et al.
Publicado: (2024)
por: Hwang, Duhun, et al.
Publicado: (2024)
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
por: Park, Jungwon, et al.
Publicado: (2024)
por: Park, Jungwon, et al.
Publicado: (2024)
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
por: Kim, Jaeill, et al.
Publicado: (2024)
por: Kim, Jaeill, et al.
Publicado: (2024)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
por: Kim, Jimyeong, et al.
Publicado: (2024)
por: Kim, Jimyeong, et al.
Publicado: (2024)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
por: Kim, Wonkyun, et al.
Publicado: (2024)
por: Kim, Wonkyun, et al.
Publicado: (2024)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
por: Kim, Jimyeong, et al.
Publicado: (2025)
por: Kim, Jimyeong, et al.
Publicado: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
por: Choi, Changin, et al.
Publicado: (2025)
por: Choi, Changin, et al.
Publicado: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
por: Ko, Jungmin, et al.
Publicado: (2026)
por: Ko, Jungmin, et al.
Publicado: (2026)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
por: Kang, Suhyun, et al.
Publicado: (2024)
por: Kang, Suhyun, et al.
Publicado: (2024)
Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation
por: Park, Jungwon, et al.
Publicado: (2026)
por: Park, Jungwon, et al.
Publicado: (2026)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
por: Song, Yeji, et al.
Publicado: (2024)
por: Song, Yeji, et al.
Publicado: (2024)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
por: Kim, Ji-Hyeon, et al.
Publicado: (2026)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
por: Kim, Jeonghyeon, et al.
Publicado: (2025)
por: Kim, Jeonghyeon, et al.
Publicado: (2025)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
por: Park, YeongHyeon, et al.
Publicado: (2024)
por: Park, YeongHyeon, et al.
Publicado: (2024)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
por: Kim, Suyoung, et al.
Publicado: (2026)
por: Kim, Suyoung, et al.
Publicado: (2026)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
por: Kim, Myunsoo, et al.
Publicado: (2025)
por: Kim, Myunsoo, et al.
Publicado: (2025)
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds
por: Kim, Keonwoo, et al.
Publicado: (2025)
por: Kim, Keonwoo, et al.
Publicado: (2025)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
por: Kim, Jihyo, et al.
Publicado: (2024)
por: Kim, Jihyo, et al.
Publicado: (2024)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
por: Lee, Uichan, et al.
Publicado: (2026)
por: Lee, Uichan, et al.
Publicado: (2026)
NEMESIS: Noise-suppressed Efficient MAE with Enhanced Superpatch Integration Strategy
por: Kim, Kyeonghun, et al.
Publicado: (2026)
por: Kim, Kyeonghun, et al.
Publicado: (2026)
Safe Semi-Supervised Contrastive Learning Using In-Distribution Data as Positive Examples
por: Kwak, Min Gu, et al.
Publicado: (2024)
por: Kwak, Min Gu, et al.
Publicado: (2024)
GuidNoise: Single-Pair Guided Diffusion for Generalized Noise Synthesis
por: Kim, Changjin, et al.
Publicado: (2025)
por: Kim, Changjin, et al.
Publicado: (2025)
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
por: Choi, Wonje, et al.
Publicado: (2024)
por: Choi, Wonje, et al.
Publicado: (2024)
PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
por: Kim, Gwanghyun, et al.
Publicado: (2024)
por: Kim, Gwanghyun, et al.
Publicado: (2024)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
por: Kim, Kangsan, et al.
Publicado: (2024)
por: Kim, Kangsan, et al.
Publicado: (2024)
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
por: Kim, Junho, et al.
Publicado: (2024)
por: Kim, Junho, et al.
Publicado: (2024)
Contrastive Learning-based Multi Modal Architecture for Emoticon Prediction by Employing Image-Text Pairs
por: Pandey, Ananya, et al.
Publicado: (2024)
por: Pandey, Ananya, et al.
Publicado: (2024)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
por: Kim, Jiyeong, et al.
Publicado: (2026)
por: Kim, Jiyeong, et al.
Publicado: (2026)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
por: Hwang, Hochul, et al.
Publicado: (2024)
por: Hwang, Hochul, et al.
Publicado: (2024)
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
por: Kim, Hyung Kyu, et al.
Publicado: (2025)
por: Kim, Hyung Kyu, et al.
Publicado: (2025)
SEAL-pose: Enhancing 3D Human Pose Estimation via a Learned Loss for Structural Consistency
por: Kim, Yeonsung, et al.
Publicado: (2026)
por: Kim, Yeonsung, et al.
Publicado: (2026)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
por: Lee, Minjae, et al.
Publicado: (2026)
por: Lee, Minjae, et al.
Publicado: (2026)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
por: Back, Kyungryul, et al.
Publicado: (2025)
por: Back, Kyungryul, et al.
Publicado: (2025)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
por: Jeong, Taejin, et al.
Publicado: (2026)
por: Jeong, Taejin, et al.
Publicado: (2026)
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
por: Lu, Mingyu, et al.
Publicado: (2025)
por: Lu, Mingyu, et al.
Publicado: (2025)
Human Interaction-Aware 3D Reconstruction from a Single Image
por: Kim, Gwanghyun, et al.
Publicado: (2026)
por: Kim, Gwanghyun, et al.
Publicado: (2026)
Frequency Composition for Compressed and Domain-Adaptive Neural Networks
por: Kwon, Yoojin, et al.
Publicado: (2025)
por: Kwon, Yoojin, et al.
Publicado: (2025)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives
por: Feng, Zhangchi, et al.
Publicado: (2024)
por: Feng, Zhangchi, et al.
Publicado: (2024)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
por: Kang, Seil, et al.
Publicado: (2025)
por: Kang, Seil, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluating Feature Attribution Methods for Electrocardiogram
por: Suh, Jangwon, et al.
Publicado: (2022) -
Towards a Better Evaluation of Out-of-Domain Generalization
por: Hwang, Duhun, et al.
Publicado: (2024) -
Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
por: Park, Jungwon, et al.
Publicado: (2024) -
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
por: Kim, Jaeill, et al.
Publicado: (2024) -
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
por: Kim, Jimyeong, et al.
Publicado: (2024)