Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Oh, Youngtaek, Ahn, Pyunghwan, Kim, Jinhyung, Song, Gwangmo, Lee, Soonyoung, Kweon, In So, Kim, Junmo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
di: Yun, Juseung, et al.
Pubblicazione: (2024)
di: Yun, Juseung, et al.
Pubblicazione: (2024)
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
di: Zhang, Chenshuang, et al.
Pubblicazione: (2024)
di: Zhang, Chenshuang, et al.
Pubblicazione: (2024)
Text-to-image Diffusion Models in Generative AI: A Survey
di: Zhang, Chenshuang, et al.
Pubblicazione: (2023)
di: Zhang, Chenshuang, et al.
Pubblicazione: (2023)
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
di: Kim, Myungchul, et al.
Pubblicazione: (2026)
di: Kim, Myungchul, et al.
Pubblicazione: (2026)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
di: Chu, Sanghyeok, et al.
Pubblicazione: (2026)
di: Chu, Sanghyeok, et al.
Pubblicazione: (2026)
Towards Understanding Dual BN In Hybrid Adversarial Training
di: Zhang, Chenshuang, et al.
Pubblicazione: (2024)
di: Zhang, Chenshuang, et al.
Pubblicazione: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
di: Zhang, Chenshuang, et al.
Pubblicazione: (2025)
di: Zhang, Chenshuang, et al.
Pubblicazione: (2025)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
di: Shin, Ukcheol, et al.
Pubblicazione: (2023)
di: Shin, Ukcheol, et al.
Pubblicazione: (2023)
ContextMix: A context-aware data augmentation method for industrial visual inspection systems
di: Kim, Hyungmin, et al.
Pubblicazione: (2024)
di: Kim, Hyungmin, et al.
Pubblicazione: (2024)
Fourier-Guided Attention Upsampling for Image Super-Resolution
di: Choi, Daejune, et al.
Pubblicazione: (2025)
di: Choi, Daejune, et al.
Pubblicazione: (2025)
The Effects of Mixed Sample Data Augmentation are Class Dependent
di: Lee, Haeil, et al.
Pubblicazione: (2023)
di: Lee, Haeil, et al.
Pubblicazione: (2023)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
di: Lee, Dongyeun, et al.
Pubblicazione: (2025)
di: Lee, Dongyeun, et al.
Pubblicazione: (2025)
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
di: Han, Gyojin, et al.
Pubblicazione: (2026)
di: Han, Gyojin, et al.
Pubblicazione: (2026)
EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision
di: Pyeon, Myeongjang, et al.
Pubblicazione: (2025)
di: Pyeon, Myeongjang, et al.
Pubblicazione: (2025)
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
di: Kim, Chris Dongjoo, et al.
Pubblicazione: (2025)
di: Kim, Chris Dongjoo, et al.
Pubblicazione: (2025)
Inspecting Explainability of Transformer Models with Additional Statistical Information
di: Nguyen, Hoang C., et al.
Pubblicazione: (2023)
di: Nguyen, Hoang C., et al.
Pubblicazione: (2023)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
di: Lee, Dong-Jae, et al.
Pubblicazione: (2026)
di: Lee, Dong-Jae, et al.
Pubblicazione: (2026)
Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis
di: Lee, Haeil, et al.
Pubblicazione: (2024)
di: Lee, Haeil, et al.
Pubblicazione: (2024)
Test-Time Mixup Augmentation for Data and Class-Specific Uncertainty Estimation in Deep Learning Image Classification
di: Lee, Hansang, et al.
Pubblicazione: (2022)
di: Lee, Hansang, et al.
Pubblicazione: (2022)
Self-supervised Transformation Learning for Equivariant Representations
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
di: Lee, Gayoung, et al.
Pubblicazione: (2025)
di: Lee, Gayoung, et al.
Pubblicazione: (2025)
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
di: Choi, Jaehyun, et al.
Pubblicazione: (2025)
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
di: Ahn, Geo, et al.
Pubblicazione: (2026)
di: Ahn, Geo, et al.
Pubblicazione: (2026)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
di: Kim, Sungnyun, et al.
Pubblicazione: (2023)
di: Kim, Sungnyun, et al.
Pubblicazione: (2023)
Learning to Explore for Stochastic Gradient MCMC
di: Kim, SeungHyun, et al.
Pubblicazione: (2024)
di: Kim, SeungHyun, et al.
Pubblicazione: (2024)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
di: Kwon, Minchan, et al.
Pubblicazione: (2026)
di: Kwon, Minchan, et al.
Pubblicazione: (2026)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
di: Choi, Jaehyun, et al.
Pubblicazione: (2024)
di: Choi, Jaehyun, et al.
Pubblicazione: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
di: Choudhury, Rohan, et al.
Pubblicazione: (2025)
di: Choudhury, Rohan, et al.
Pubblicazione: (2025)
Identifiable Token Correspondence for World Models
di: Kim, Youngin, et al.
Pubblicazione: (2026)
di: Kim, Youngin, et al.
Pubblicazione: (2026)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
di: Li, Can, et al.
Pubblicazione: (2025)
di: Li, Can, et al.
Pubblicazione: (2025)
360 in the Wild: Dataset for Depth Prediction and View Synthesis
di: Park, Kibaek, et al.
Pubblicazione: (2024)
di: Park, Kibaek, et al.
Pubblicazione: (2024)
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
di: Kim, Sumin, et al.
Pubblicazione: (2026)
di: Kim, Sumin, et al.
Pubblicazione: (2026)
Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
di: Oh, Youngmin, et al.
Pubblicazione: (2026)
di: Oh, Youngmin, et al.
Pubblicazione: (2026)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models
di: Hwang, Jisung, et al.
Pubblicazione: (2025)
di: Hwang, Jisung, et al.
Pubblicazione: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
VisioFirm: Cross-Platform AI-assisted Annotation Tool for Computer Vision
di: Ghazouali, Safouane El, et al.
Pubblicazione: (2025)
di: Ghazouali, Safouane El, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
di: Oh, Youngtaek, et al.
Pubblicazione: (2024) -
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
di: Yun, Juseung, et al.
Pubblicazione: (2024) -
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
di: Zhang, Chenshuang, et al.
Pubblicazione: (2024) -
Text-to-image Diffusion Models in Generative AI: A Survey
di: Zhang, Chenshuang, et al.
Pubblicazione: (2023) -
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
di: Kim, Myungchul, et al.
Pubblicazione: (2026)