LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Hashemi, Mohammad Abuzar, Li, Zhanghexuan, Chauhan, Mihir, Shen, Yan, Satbhai, Abhishek, Ali, Mir Basheer, Gao, Mingchen, Srihari, Sargur |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Learning Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024)
by: Chauhan, Mihir, et al.
Published: (2024)
Vision-Language Model Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024)
by: Chauhan, Mihir, et al.
Published: (2024)
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
by: Mittal, Prateek, et al.
Published: (2023)
by: Mittal, Prateek, et al.
Published: (2023)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
by: Hossain, Eftekhar, et al.
Published: (2024)
by: Hossain, Eftekhar, et al.
Published: (2024)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
by: Chen, Junan, et al.
Published: (2025)
by: Chen, Junan, et al.
Published: (2025)
Current Symmetry Group Equivariant Convolution Frameworks for Representation Learning
by: Basheer, Ramzan, et al.
Published: (2024)
by: Basheer, Ramzan, et al.
Published: (2024)
Figuring out Figures: Using Textual References to Caption Scientific Figures
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine
by: Huang, Jiatan, et al.
Published: (2024)
by: Huang, Jiatan, et al.
Published: (2024)
Captioning Visualizations with Large Language Models (CVLLM): A Tutorial
by: Carenini, Giuseppe, et al.
Published: (2024)
by: Carenini, Giuseppe, et al.
Published: (2024)
Aligned Textual Scoring Rules
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Manipulating a Tetris-Inspired 3D Video Representation
by: Godbole, Mihir
Published: (2024)
by: Godbole, Mihir
Published: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Information Extraction: An application to the domain of hyper-local financial data on developing countries
by: Royesh, Abuzar, et al.
Published: (2024)
by: Royesh, Abuzar, et al.
Published: (2024)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
TopoAlign: Topology-Aware Visual Representation Alignment
by: Yan, Xinyuan, et al.
Published: (2026)
by: Yan, Xinyuan, et al.
Published: (2026)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
by: Chen, Danlu, et al.
Published: (2024)
by: Chen, Danlu, et al.
Published: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
by: Liang, Ziqi, et al.
Published: (2024)
by: Liang, Ziqi, et al.
Published: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
by: Bandraupalli, Srihari, et al.
Published: (2025)
by: Bandraupalli, Srihari, et al.
Published: (2025)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Instance-Aligned Captions for Explainable Video Anomaly Detection
by: Song, Inpyo, et al.
Published: (2026)
by: Song, Inpyo, et al.
Published: (2026)
Aligning Actions and Walking to LLM-Generated Textual Descriptions
by: Chivereanu, Radu, et al.
Published: (2024)
by: Chivereanu, Radu, et al.
Published: (2024)
Towards Aligning Language Models with Textual Feedback
by: Lloret, Saüc Abadal, et al.
Published: (2024)
by: Lloret, Saüc Abadal, et al.
Published: (2024)
Test-Time Conditioning with Representation-Aligned Visual Features
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
Continual Domain Adversarial Adaptation via Double-Head Discriminators
by: Shen, Yan, et al.
Published: (2024)
by: Shen, Yan, et al.
Published: (2024)
CLEAR: Character Unlearning in Textual and Visual Modalities
by: Dontsov, Alexey, et al.
Published: (2024)
by: Dontsov, Alexey, et al.
Published: (2024)
reslife: Residual Lifetime Analysis Tool in R
by: Wang, Zekai, et al.
Published: (2023)
by: Wang, Zekai, et al.
Published: (2023)
ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
by: Madani, Navid, et al.
Published: (2025)
by: Madani, Navid, et al.
Published: (2025)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
by: Anaissi, Ali, et al.
Published: (2025)
by: Anaissi, Ali, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
by: Fadeev, Egor, et al.
Published: (2025)
by: Fadeev, Egor, et al.
Published: (2025)
The Automated Verification of Textual Claims (AVeriTeC) Shared Task
by: Schlichtkrull, Michael, et al.
Published: (2024)
by: Schlichtkrull, Michael, et al.
Published: (2024)
Beyond the Textual: Generating Coherent Visual Options for MCQs
by: Wang, Wanqiang, et al.
Published: (2025)
by: Wang, Wanqiang, et al.
Published: (2025)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
by: Sukhani, Siddhant, et al.
Published: (2025)
by: Sukhani, Siddhant, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
AutoVisual Fusion Suite: A Comprehensive Evaluation of Image Segmentation and Voice Conversion Tools on HuggingFace Platform
by: Hashemi, Amirreza
Published: (2023)
by: Hashemi, Amirreza
Published: (2023)
A Comprehensive Review of Visual-Textual Sentiment Analysis from Social Media Networks
by: Al-Tameemi, Israa Khalaf Salman, et al.
Published: (2022)
by: Al-Tameemi, Israa Khalaf Salman, et al.
Published: (2022)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
by: Lee, Yujian, et al.
Published: (2026)
by: Lee, Yujian, et al.
Published: (2026)
Similar Items
-
Self-Supervised Learning Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024) -
Vision-Language Model Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024) -
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
by: Mittal, Prateek, et al.
Published: (2023) -
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024) -
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
by: Hossain, Eftekhar, et al.
Published: (2024)