ARGENT: Adaptive Hierarchical Image-Text Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Huynh, Chuong, Souri, Hossein, Kumar, Abhinav, Petsiuk, Vitali, Mohan, Deen Dayal, Kumar, Suren |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
by: Mohan, Deen Dayal, et al.
Published: (2026)
by: Mohan, Deen Dayal, et al.
Published: (2026)
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
by: Min, Juhong, et al.
Published: (2026)
by: Min, Juhong, et al.
Published: (2026)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
by: Petsiuk, Vitali, et al.
Published: (2024)
by: Petsiuk, Vitali, et al.
Published: (2024)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
All-in-One Conditioning for Text-to-Image Synthesis
by: Jayasekara, Hirunima, et al.
Published: (2026)
by: Jayasekara, Hirunima, et al.
Published: (2026)
Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
SH-SAS: An Implicit Neural Representation for Complex Spherical-Harmonic Scattering Fields for 3D Synthetic Aperture Sonar
by: Vengurlekar, Omkar Shailendra, et al.
Published: (2025)
by: Vengurlekar, Omkar Shailendra, et al.
Published: (2025)
Adaptive Object Detection for Indoor Navigation Assistance: A Performance Evaluation of Real-Time Algorithms
by: Pratap, Abhinav, et al.
Published: (2025)
by: Pratap, Abhinav, et al.
Published: (2025)
Hyperbolic Image-Text Representations
by: Desai, Karan, et al.
Published: (2023)
by: Desai, Karan, et al.
Published: (2023)
Improved Probabilistic Image-Text Representations
by: Chun, Sanghyuk
Published: (2023)
by: Chun, Sanghyuk
Published: (2023)
Efficient and High-Fidelity Omni Modality Retrieval
by: Huynh, Chuong, et al.
Published: (2026)
by: Huynh, Chuong, et al.
Published: (2026)
Finetuning-Free Personalization of Text to Image Generation via Hypernetworks
by: Shrestha, Sagar, et al.
Published: (2025)
by: Shrestha, Sagar, et al.
Published: (2025)
A Fast and Efficient Modern BERT based Text-Conditioned Diffusion Model for Medical Image Segmentation
by: Dhara, Venkata Siddharth, et al.
Published: (2025)
by: Dhara, Venkata Siddharth, et al.
Published: (2025)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025)
by: Das, Swadhin, et al.
Published: (2025)
CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text
by: Rani, Anju, et al.
Published: (2025)
by: Rani, Anju, et al.
Published: (2025)
Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation
by: Zare, Mohammad
Published: (2025)
by: Zare, Mohammad
Published: (2025)
Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning
by: Role, François, et al.
Published: (2025)
by: Role, François, et al.
Published: (2025)
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
by: Kumar, Amandeep, et al.
Published: (2026)
by: Kumar, Amandeep, et al.
Published: (2026)
Model-Agnostic Gender Bias Control for Text-to-Image Generation via Sparse Autoencoder
by: Wu, Chao, et al.
Published: (2025)
by: Wu, Chao, et al.
Published: (2025)
Turbulence Strength $C_n^2$ Estimation from Video using Physics-based Deep Learning
by: Saha, Ripon Kumar, et al.
Published: (2024)
by: Saha, Ripon Kumar, et al.
Published: (2024)
Text-to-Image GAN with Pretrained Representations
by: You, Xiaozhou, et al.
Published: (2024)
by: You, Xiaozhou, et al.
Published: (2024)
Hybrid CNN with Chebyshev Polynomial Expansion for Medical Image Analysis
by: Roy, Abhinav, et al.
Published: (2025)
by: Roy, Abhinav, et al.
Published: (2025)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
by: Jayatilaka, Gihan, et al.
Published: (2025)
by: Jayatilaka, Gihan, et al.
Published: (2025)
ADROIT: A Self-Supervised Framework for Learning Robust Representations for Active Learning
by: Banerjee, Soumya, et al.
Published: (2025)
by: Banerjee, Soumya, et al.
Published: (2025)
WriteViT: Handwritten Text Generation with Vision Transformer
by: Nam, Dang Hoai, et al.
Published: (2025)
by: Nam, Dang Hoai, et al.
Published: (2025)
HTR-ConvText: Leveraging Convolution and Textual Information for Handwritten Text Recognition
by: Truc, Pham Thach Thanh, et al.
Published: (2025)
by: Truc, Pham Thach Thanh, et al.
Published: (2025)
Continual Learning: Less Forgetting, More OOD Generalization via Adaptive Contrastive Replay
by: Rezaei, Hossein, et al.
Published: (2024)
by: Rezaei, Hossein, et al.
Published: (2024)
A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
by: Singh, Nayan Kumar
Published: (2025)
by: Singh, Nayan Kumar
Published: (2025)
CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
by: Mohan, Riya, et al.
Published: (2026)
by: Mohan, Riya, et al.
Published: (2026)
Boosting Open Set Recognition Performance through Modulated Representation Learning
by: Kundu, Amit Kumar, et al.
Published: (2025)
by: Kundu, Amit Kumar, et al.
Published: (2025)
Bayesian Multi-Scale Neural Network for Crowd Counting
by: Sagar, Abhinav
Published: (2020)
by: Sagar, Abhinav
Published: (2020)
ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy Images
by: Bourriez, Nicolas, et al.
Published: (2023)
by: Bourriez, Nicolas, et al.
Published: (2023)
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
by: Das, Aryan, et al.
Published: (2026)
by: Das, Aryan, et al.
Published: (2026)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
Interpretable Image Emotion Recognition: A Domain Adaptation Approach Using Facial Expressions
by: Kumar, Puneet, et al.
Published: (2020)
by: Kumar, Puneet, et al.
Published: (2020)
Adaptive Hierarchical Certification for Segmentation using Randomized Smoothing
by: Anani, Alaa, et al.
Published: (2024)
by: Anani, Alaa, et al.
Published: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation
by: Abdelsamad, Mohamed, et al.
Published: (2025)
by: Abdelsamad, Mohamed, et al.
Published: (2025)
Similar Items
-
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
by: Mohan, Deen Dayal, et al.
Published: (2026) -
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
by: Min, Juhong, et al.
Published: (2026) -
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
by: Petsiuk, Vitali, et al.
Published: (2024) -
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024) -
All-in-One Conditioning for Text-to-Image Synthesis
by: Jayasekara, Hirunima, et al.
Published: (2026)