Veagle: Advancements in Multimodal Representation Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Chawla, Rajat, Datta, Arkajit, Verma, Tushar, Jha, Adarsh, Gautam, Anmol, Vatsal, Ayush, Chaterjee, Sukrit, NS, Mukunda, Bhola, Ishaan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
di: Rahman, Abdur, et al.
Pubblicazione: (2024)
di: Rahman, Abdur, et al.
Pubblicazione: (2024)
AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
di: Datta, Arkajit, et al.
Pubblicazione: (2024)
di: Datta, Arkajit, et al.
Pubblicazione: (2024)
GUIDE: Graphical User Interface Data for Execution
di: Chawla, Rajat, et al.
Pubblicazione: (2024)
di: Chawla, Rajat, et al.
Pubblicazione: (2024)
SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer
di: Gautam, Anmol, et al.
Pubblicazione: (2024)
di: Gautam, Anmol, et al.
Pubblicazione: (2024)
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
di: Bhola, Ishaan, et al.
Pubblicazione: (2025)
di: Bhola, Ishaan, et al.
Pubblicazione: (2025)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
di: Qian, Fan, et al.
Pubblicazione: (2024)
di: Qian, Fan, et al.
Pubblicazione: (2024)
Interactive Video Generation via Domain Adaptation
di: Rawal, Ishaan, et al.
Pubblicazione: (2025)
di: Rawal, Ishaan, et al.
Pubblicazione: (2025)
Principled Multimodal Representation Learning
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Television Discourse Decoded: Comprehensive Multimodal Analytics at Scale
di: Agarwal, Anmol, et al.
Pubblicazione: (2024)
di: Agarwal, Anmol, et al.
Pubblicazione: (2024)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
di: Liu, Peipei, et al.
Pubblicazione: (2023)
di: Liu, Peipei, et al.
Pubblicazione: (2023)
Calibrated Multimodal Representation Learning with Missing Modalities
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
di: Liu, Xiaohao, et al.
Pubblicazione: (2025)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
di: Wang, Bing, et al.
Pubblicazione: (2025)
di: Wang, Bing, et al.
Pubblicazione: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
di: Shah, Siddhant Bikram, et al.
Pubblicazione: (2024)
di: Shah, Siddhant Bikram, et al.
Pubblicazione: (2024)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
di: Lu, Renjie, et al.
Pubblicazione: (2026)
di: Lu, Renjie, et al.
Pubblicazione: (2026)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
di: Wang, Bokang, et al.
Pubblicazione: (2026)
di: Wang, Bokang, et al.
Pubblicazione: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
di: Zhao, Ruixiang, et al.
Pubblicazione: (2024)
di: Zhao, Ruixiang, et al.
Pubblicazione: (2024)
Detached and Interactive Multimodal Learning
di: Fan, Yunfeng, et al.
Pubblicazione: (2024)
di: Fan, Yunfeng, et al.
Pubblicazione: (2024)
Multimodal Sentiment Analysis Based on Causal Reasoning
di: Chen, Fuhai, et al.
Pubblicazione: (2024)
di: Chen, Fuhai, et al.
Pubblicazione: (2024)
Retrieval-Augmented Multimodal Model for Fake News Detection
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
Learning Brain Representation with Hierarchical Visual Embeddings
di: Zheng, Jiawen, et al.
Pubblicazione: (2026)
di: Zheng, Jiawen, et al.
Pubblicazione: (2026)
Learning Video Context as Interleaved Multimodal Sequences
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
di: Chen, Sen, et al.
Pubblicazione: (2022)
di: Chen, Sen, et al.
Pubblicazione: (2022)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
di: Zhang, Lei, et al.
Pubblicazione: (2025)
di: Zhang, Lei, et al.
Pubblicazione: (2025)
Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
di: Chen, Xiaolin, et al.
Pubblicazione: (2025)
di: Chen, Xiaolin, et al.
Pubblicazione: (2025)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
di: Quan, Weize, et al.
Pubblicazione: (2024)
di: Quan, Weize, et al.
Pubblicazione: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
di: Cheng, Zhi-Qi, et al.
Pubblicazione: (2024)
di: Cheng, Zhi-Qi, et al.
Pubblicazione: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
Self-supervised Photographic Image Layout Representation Learning
di: Zhao, Zhaoran, et al.
Pubblicazione: (2024)
di: Zhao, Zhaoran, et al.
Pubblicazione: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion
di: Zhao, Yu, et al.
Pubblicazione: (2024)
di: Zhao, Yu, et al.
Pubblicazione: (2024)
Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis
di: Liu, Hao, et al.
Pubblicazione: (2025)
di: Liu, Hao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
di: Rahman, Abdur, et al.
Pubblicazione: (2024) -
AUTONODE: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
di: Datta, Arkajit, et al.
Pubblicazione: (2024) -
GUIDE: Graphical User Interface Data for Execution
di: Chawla, Rajat, et al.
Pubblicazione: (2024) -
SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer
di: Gautam, Anmol, et al.
Pubblicazione: (2024) -
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
di: Bhola, Ishaan, et al.
Pubblicazione: (2025)