LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hashemi, Mohammad Abuzar, Li, Zhanghexuan, Chauhan, Mihir, Shen, Yan, Satbhai, Abhishek, Ali, Mir Basheer, Gao, Mingchen, Srihari, Sargur |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2021
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Supervised Learning Based Handwriting Verification
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
Vision-Language Model Based Handwriting Verification
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
di: Chauhan, Mihir, et al.
Pubblicazione: (2024)
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
di: Mittal, Prateek, et al.
Pubblicazione: (2023)
di: Mittal, Prateek, et al.
Pubblicazione: (2023)
OSCaR: Object State Captioning and State Change Representation
di: Nguyen, Nguyen, et al.
Pubblicazione: (2024)
di: Nguyen, Nguyen, et al.
Pubblicazione: (2024)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
di: Hossain, Eftekhar, et al.
Pubblicazione: (2024)
di: Hossain, Eftekhar, et al.
Pubblicazione: (2024)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
di: Chen, Junan, et al.
Pubblicazione: (2025)
di: Chen, Junan, et al.
Pubblicazione: (2025)
Current Symmetry Group Equivariant Convolution Frameworks for Representation Learning
di: Basheer, Ramzan, et al.
Pubblicazione: (2024)
di: Basheer, Ramzan, et al.
Pubblicazione: (2024)
Figuring out Figures: Using Textual References to Caption Scientific Figures
di: Cao, Stanley, et al.
Pubblicazione: (2024)
di: Cao, Stanley, et al.
Pubblicazione: (2024)
RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine
di: Huang, Jiatan, et al.
Pubblicazione: (2024)
di: Huang, Jiatan, et al.
Pubblicazione: (2024)
Captioning Visualizations with Large Language Models (CVLLM): A Tutorial
di: Carenini, Giuseppe, et al.
Pubblicazione: (2024)
di: Carenini, Giuseppe, et al.
Pubblicazione: (2024)
Aligned Textual Scoring Rules
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
di: Park, Kyu Ri, et al.
Pubblicazione: (2025)
di: Park, Kyu Ri, et al.
Pubblicazione: (2025)
Manipulating a Tetris-Inspired 3D Video Representation
di: Godbole, Mihir
Pubblicazione: (2024)
di: Godbole, Mihir
Pubblicazione: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
Information Extraction: An application to the domain of hyper-local financial data on developing countries
di: Royesh, Abuzar, et al.
Pubblicazione: (2024)
di: Royesh, Abuzar, et al.
Pubblicazione: (2024)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
di: Dong, Sixun, et al.
Pubblicazione: (2025)
di: Dong, Sixun, et al.
Pubblicazione: (2025)
TopoAlign: Topology-Aware Visual Representation Alignment
di: Yan, Xinyuan, et al.
Pubblicazione: (2026)
di: Yan, Xinyuan, et al.
Pubblicazione: (2026)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
di: Chen, Danlu, et al.
Pubblicazione: (2024)
di: Chen, Danlu, et al.
Pubblicazione: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
di: Sun, Zhengyang, et al.
Pubblicazione: (2026)
di: Sun, Zhengyang, et al.
Pubblicazione: (2026)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
From Image Captioning to Visual Storytelling
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
Instance-Aligned Captions for Explainable Video Anomaly Detection
di: Song, Inpyo, et al.
Pubblicazione: (2026)
di: Song, Inpyo, et al.
Pubblicazione: (2026)
Aligning Actions and Walking to LLM-Generated Textual Descriptions
di: Chivereanu, Radu, et al.
Pubblicazione: (2024)
di: Chivereanu, Radu, et al.
Pubblicazione: (2024)
Towards Aligning Language Models with Textual Feedback
di: Lloret, Saüc Abadal, et al.
Pubblicazione: (2024)
di: Lloret, Saüc Abadal, et al.
Pubblicazione: (2024)
Test-Time Conditioning with Representation-Aligned Visual Features
di: Sereyjol-Garros, Nicolas, et al.
Pubblicazione: (2026)
di: Sereyjol-Garros, Nicolas, et al.
Pubblicazione: (2026)
Continual Domain Adversarial Adaptation via Double-Head Discriminators
di: Shen, Yan, et al.
Pubblicazione: (2024)
di: Shen, Yan, et al.
Pubblicazione: (2024)
CLEAR: Character Unlearning in Textual and Visual Modalities
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
reslife: Residual Lifetime Analysis Tool in R
di: Wang, Zekai, et al.
Pubblicazione: (2023)
di: Wang, Zekai, et al.
Pubblicazione: (2023)
ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
di: Madani, Navid, et al.
Pubblicazione: (2025)
di: Madani, Navid, et al.
Pubblicazione: (2025)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
di: Anaissi, Ali, et al.
Pubblicazione: (2025)
di: Anaissi, Ali, et al.
Pubblicazione: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
di: Li, Yuying, et al.
Pubblicazione: (2025)
di: Li, Yuying, et al.
Pubblicazione: (2025)
LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
di: Fadeev, Egor, et al.
Pubblicazione: (2025)
di: Fadeev, Egor, et al.
Pubblicazione: (2025)
The Automated Verification of Textual Claims (AVeriTeC) Shared Task
di: Schlichtkrull, Michael, et al.
Pubblicazione: (2024)
di: Schlichtkrull, Michael, et al.
Pubblicazione: (2024)
Beyond the Textual: Generating Coherent Visual Options for MCQs
di: Wang, Wanqiang, et al.
Pubblicazione: (2025)
di: Wang, Wanqiang, et al.
Pubblicazione: (2025)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
di: Sukhani, Siddhant, et al.
Pubblicazione: (2025)
di: Sukhani, Siddhant, et al.
Pubblicazione: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
AutoVisual Fusion Suite: A Comprehensive Evaluation of Image Segmentation and Voice Conversion Tools on HuggingFace Platform
di: Hashemi, Amirreza
Pubblicazione: (2023)
di: Hashemi, Amirreza
Pubblicazione: (2023)
A Comprehensive Review of Visual-Textual Sentiment Analysis from Social Media Networks
di: Al-Tameemi, Israa Khalaf Salman, et al.
Pubblicazione: (2022)
di: Al-Tameemi, Israa Khalaf Salman, et al.
Pubblicazione: (2022)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
di: Lee, Yujian, et al.
Pubblicazione: (2026)
di: Lee, Yujian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Self-Supervised Learning Based Handwriting Verification
di: Chauhan, Mihir, et al.
Pubblicazione: (2024) -
Vision-Language Model Based Handwriting Verification
di: Chauhan, Mihir, et al.
Pubblicazione: (2024) -
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
di: Mittal, Prateek, et al.
Pubblicazione: (2023) -
OSCaR: Object State Captioning and State Change Representation
di: Nguyen, Nguyen, et al.
Pubblicazione: (2024) -
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
di: Hossain, Eftekhar, et al.
Pubblicazione: (2024)